Model or dataset
jeinlee1991/chinese-llm-benchmark avatar
jeinlee1991/chinese-llm-benchmark

jeinlee1991/chinese-llm-benchmark: what the ReLE leaderboard measures and what it does not

非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat等商用模型, 以及step3.5-flash、kimi-k2.6、ernie4.5、MiniMax-M2.7、deepseek-v4、Qwen3.6、llama4、智谱GLM-5.1、MiMo-V2、LongCat、gemma4、mistral等开源大模型。不仅提供排行榜,也提供规模超200万的大模型缺陷库!方便广大社区研究分析、改进大模型。

6,457 stars264 forksUnknownLicense varies

At a glance

What is it?
ReLE is a Chinese-language capability benchmark that scores hundreds of commercial and open models across seven domains and roughly 300 sub-dimensions, and pairs the rankings with a defect library. It is a reference table, not a test harness you run yourself.
Who is it for?
Use chinese-llm-benchmark when you need a Chinese-language capability signal across many models at once, especially for domains like law, medicine or education that English leaderboards do not cover. Do not use it as your only selection input, and do not treat it as a runnable harness: the repository's visible top level holds leaderboard, opendata and eval directories, but the README documents no local install or scoring command.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
GitHub does not report a main language for this repository.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

Which Chinese LLM is best is the wrong question for this repository

The README opens with a claim that no single model wins everything, and the structure backs that up. Instead of one aggregate score, ReLE splits evaluation into seven domains: education, healthcare and mental health, finance, law and public administration, reasoning and mathematical calculation, language and instruction following, and agent and tool calling. Under those sit roughly 300 finer dimensions, and the README names dental medicine and high school Chinese as examples. That granularity is the actual product. A team building a medical triage assistant and a team building a contract review tool get different useful rows from the same table.

The second half of the pitch is unusual. The project states it ships a defect library with more than two million entries, described as material for the community to analyse and improve models. A leaderboard tells you who scored higher; a defect library tells you how a model failed. For anyone doing error analysis before fine-tuning, the second artifact is the more interesting one, though the README does not describe its schema or how to query it.

The audience is narrower than the model count suggests. This is for people choosing a Chinese-capable model for a specific vertical, and for researchers who want a cross-model failure corpus. It is not for someone who wants to run the benchmark against their own private model locally, unless they go through the free evaluation service the README offers to private model owners via a WeChat contact.

How ReLE is structured: domains, sub-dimensions and a defect library

ReLE stands for Really Reliable Live Evaluation, and the README notes the project was previously called CLiB. The word live matters: the update log shows near-weekly version bumps, with models added and older ones deleted. In the 2026-09-08 v5.11.7 entry, gpt-6-astra is added and a long list of stale models is removed, including gpt-5.1, claude-sonnet-4.5, DeepSeek-V3.1 and several Qwen3 variants. That pruning policy is a design decision with consequences. The board stays current, but a score you cited six months ago may no longer exist in the table.

The seven domains are further divided into exam-style and task-style sections. Education runs from primary school subjects through the gaokao, with middle school, graduate entrance and teacher certification listed. Finance splits into accounting, banking, insurance, securities, other financial qualifications, financial fundamentals and financial applications. Reasoning covers deductive reasoning, common-sense reasoning, BBH symbolic reasoning, arithmetic, table question answering, table summarisation, olympiad maths at three levels, sudoku and more. Language and instruction following covers idiom understanding, sentiment analysis, textual entailment, classification, information extraction, reading comprehension, pronoun resolution, classical poetry matching, Chinese instruction following and character glyph tasks. Agent and tool calling is measured through TAU and BFCL-V3. Coding is measured through livecodebench and Terminal-Bench-2.0. There is also a section that integrates LMArena and AA scores rather than producing its own.

A separate multimodal README covers multimodal understanding and image generation, and a dedicated file covers image generation evaluation. The repository top level also contains leaderboard, opendata, eval, docs and github_star_data directories, plus a weekly new-models file. The presence of an eval directory suggests evaluation code exists, but the README does not document how to invoke it, and I have not run it.

Installing chinese-llm-benchmark: there is nothing to install

This is the part that trips people up. The README is a leaderboard document, not a getting-started guide. It lists no pip install, no npm package, no Docker image, no CLI entry point and no environment variables. There is no configuration snippet to copy. If you arrived from a search for a Chinese LLM benchmark you can run on your own machine, this repository is not that, at least not from what the README documents.

The practical first use is to read the tables. Clone the repository and open the leaderboard directory, or read the rendered README on the default branch, which is main. The README also points to a homepage at nonelinear.com, which is where the project presents itself outside GitHub.

bash
git clone https://github.com/jeinlee1991/chinese-llm-benchmark
cd chinese-llm-benchmark
ls leaderboard opendata eval docs

The listing is the useful output. You should see the leaderboard tables, the open data directory and the eval directory. What you will not find is a documented command that reproduces a score.

If you want your own private model evaluated, the README's route is not code. It offers free evaluation for private models and directs you to contact the NoneLinear ReLE benchmark team on WeChat. That is a service relationship, not a self-serve pipeline, and it means your model leaves your control to be scored. For teams under data-handling constraints, that is a real gate.

For citation, the README provides a "Cite Us" section and points to the technical report, ReLE: A Scalable System and Structured Benchmark for Diagnosing Capability Anisotropy in Chinese LLMs, on arXiv. If you need to know how a score was produced, that paper is the place to look, not the repository.

The TODO markers are the honest part and the weak part

Scan the table of contents and you find TODO attached to specific dimensions: middle school exams, higher education, graduate entrance exams, teacher certification, middle school olympiad, map reasoning, spatial reasoning, amount capitalisation conversion, date calculation, Pinyin, typo detection, sentence understanding, punctuation and simplified-traditional character conversion. Some of these are marked TODO in the index while the surrounding sections are populated.

Read that as a coverage map, not a defect. Publishing the gaps is better than implying uniform coverage. But it means the domain count of seven overstates what is actually measured today. If your use case depends on punctuation handling or traditional-to-simplified conversion, the leaderboard may not have a row for you yet, and the model that looks best on the aggregate view is not necessarily the one that handles your specific task.

The second limitation is the deletion policy. Removing stale models keeps the board readable, but it also means the benchmark is not a stable historical record. If you are comparing a model you deployed last year against a current one, the old row may be gone. The CHANGELOG.md file at the repository root is the place to check what was removed and when, since the README's update log only summarises recent versions.

A third constraint: the README does not state the licence. That matters if you intend to reuse the defect library or the leaderboard data in a product. The repository has no licence field documented, so treat redistribution as unresolved until you confirm it with the maintainers.

ReLE against C-Eval and the Open LLM Leaderboard

The obvious comparison is C-Eval, the Chinese evaluation suite that appears in search interest around this topic. C-Eval is a fixed question set with a published harness: you download it, point it at a model, and get scores you can reproduce. ReLE is the opposite shape. It is a continuously updated leaderboard run by a team, with results published as tables, plus a defect corpus. If you need to score your own checkpoint tonight, C-Eval is the tool. If you need to know which of several hundred hosted Chinese models handles dental questions better, ReLE is the tool.

The second comparison is the Open LLM Leaderboard on Hugging Face. That board is model-submission oriented and centred on English academic tasks. ReLE's centre of gravity is Chinese professional and exam domains, and it explicitly integrates LMArena and AA scores in a separate section rather than competing with them. The difference in approach is who runs the evaluation: community submissions and standardised harnesses on one side, a maintained internal pipeline with a published technical report on the other.

The trade-off is transparency for coverage. A fixed harness gives you reproducibility and lets you audit every question. A maintained leaderboard gives you breadth across 398 models and domains that no single academic suite covers, at the cost of trusting the team's methodology. The arXiv report is what makes that trust checkable, and it is the reason the report matters more than the tables.

Maintenance cadence, versioning and what it costs you

The last push to the repository was on 2026-09-08, and the README's update log shows versions landing roughly weekly through 2026, with v5.11.7 on 2026-09-08, v5.11.6 on 2026-09-04, v5.11.5 on 2026-09-02 and v5.11.4 on 2026-08-29. Recent releases listed are v5.10 from 2026-04-21, v5.9 from 2026-04-18 and v5.8.23 from 2026-04-15. The cadence is the project's main asset and its main maintenance cost for you.

That cost is real. If you cite a ReLE score in an internal decision document, it has a shelf life measured in weeks. Pinning to a version number, the way the README labels each entry, is the only way to make a citation stable, and even then the underlying model may be deprecated by its vendor. The README's own deletion lists show vendors retiring model snapshots constantly.

On licence, the README does not state one. It offers free evaluation for private models, and the defect library is described as available for community research. Whether that extends to commercial reuse of the data is not addressed. Do not assume permissive terms; ask the maintainers before you build on the dataset.

Editorial conclusion

Use chinese-llm-benchmark when you need a Chinese-language capability signal across many models at once, especially for domains like law, medicine or education that English leaderboards do not cover. Do not use it as your only selection input, and do not treat it as a runnable harness: the repository's visible top level holds leaderboard, opendata and eval directories, but the README documents no local install or scoring command. Before you commit, open the arXiv technical report and check whether the sub-dimension you care about is marked TODO, because several listed dimensions are still placeholders.

Frequently asked questions

Which Chinese LLM is the best according to chinese-llm-benchmark?

The README states there is no all-round winner, and the leaderboard is deliberately split across seven domains and roughly 300 sub-dimensions so that different models lead in different areas. That is the point of the design rather than a gap in it.

What are the main LLM benchmarks, and where does chinese-llm-benchmark fit?

The repository positions itself alongside LMArena and AA, whose scores it integrates in a separate section, and it covers Chinese professional and exam domains that English-centric suites do not. It also points to its own technical report on arXiv for the methodology.

How does LLM benchmarking work in chinese-llm-benchmark?

The project runs its own evaluation pipeline and publishes results as leaderboard tables, with a defect library of more than two million entries alongside them. The README does not document a local harness, so the mechanism is described in the arXiv technical report rather than in installable code.

Official sources

  1. Issues
  2. jeinlee1991/chinese-llm-benchmark on GitHub
  3. Project website
  4. README
  5. Releases
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/jeinlee1991-chinese-llm-benchmark.svg)](https://hysenlabs.com/projects/jeinlee1991-chinese-llm-benchmark)
Community notes

Community notes