Model or dataset
jeinlee1991/chinese-llm-benchmark avatar
jeinlee1991/chinese-llm-benchmark

ReLE (chinese-llm-benchmark): A Chinese LLM Leaderboard and a 2 Million Item Defect Library

非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat等商用模型, 以及step3.5-flash、kimi-k2.6、ernie4.5、MiniMax-M2.7、deepseek-v4、Qwen3.6、llama4、智谱GLM-5.1、MiMo-V2、LongCat、gemma4、mistral等开源大模型。不仅提供排行榜,也提供规模超200万的大模型缺陷库!方便广大社区研究分析、改进大模型。

6,442 stars265 forksUnknownLicense varies

At a glance

What is it?
ReLE is a continuously updated Chinese-language LLM evaluation project that publishes leaderboards across seven domains and roughly 300 sub-dimensions, alongside what the README describes as a defect library of more than 2 million items. It is a data and reporting project, not a harness you install.
Who is it for?
Adopt ReLE if you need Chinese-language capability comparisons across education, medicine, finance, law, reasoning, instruction following and agent tool calling, and you accept that results arrive on the maintainer's schedule. Do not adopt it if you need a harness you can run against your own prompts, or if your evaluation targets English or non-Chinese-language workloads.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
GitHub does not report a main language for this repository.

Answers come from the project's GitHub data, last synced on September 15, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What ReLE Actually Is, and Who It Is Built For

ReLE stands for Really Reliable Live Evaluation for LLM, and the README notes it was formerly called CLiB. The project is a benchmark and reporting effort rather than a library. Its output is a set of leaderboards plus a defect corpus, and the README states that the defect library exceeds 2 million entries. That framing matters: if you are looking for a package to add to a test suite, this is not it. If you are choosing a Chinese-language model for a product, or you want to know where a model breaks rather than only how often it succeeds, the two artifacts map onto two different questions. The stated audience is broad. The README describes coverage of 398 models at the time of writing, spanning commercial systems such as chatgpt, gpt-6, gemini-3.1-pro, Claude-5, grok-4.6, ERNIE-X1.1 and qwen3.8-max, alongside open weight families including llama4, deepseek-v4, Qwen3.8, GLM-5.3, gemma4 and mistral. The README also offers free evaluation for private models through a WeChat contact, which signals that the intended users include teams with in-house models that cannot be shipped to a public leaderboard.

Seven Domains, Roughly 300 Sub-Dimensions, and a Defect Library

The evaluation structure is the most distinctive part of the project. The README lists seven top-level domains: education, healthcare and mental health, finance, law and public administration, reasoning and mathematical calculation, language and instruction following, and agent and tool calling. Under those sit roughly 300 finer dimensions, and the README gives dental care and high school Chinese language as examples. That granularity is the point. A single aggregate score hides the fact that a model may handle high school mathematics while failing at pharmacy questions or at converting Chinese character forms. The README also links a technical report titled ReLE: A Scalable System and Structured Benchmark for Diagnosing Capability Anisotropy in Chinese LLMs, which is where the methodology presumably lives. The defect library is the second artifact, and the README positions it for community research and model improvement rather than for procurement decisions. Note that a defect corpus of that size is only useful if the defects are labeled and reproducible. The README does not describe the labeling scheme or the format, so treat the count as a headline until you inspect the data itself.

How the Leaderboards Are Organized

The table of contents reveals a layered structure rather than one ranking. There is a multimodal section with separate multimodal understanding and image generation leaderboards. The general capability section splits into reasoning models, commercial models including paid APIs for open weight models, and open source models. Domain sections then break down further: education runs from primary school subjects through middle school, high school, and the gaokao, with several categories marked TODO. Healthcare splits into physician, nursing, pharmacist, medical technology, basic medical knowledge, medical postgraduate exams, and mental health. Finance covers accounting, banking, insurance, securities, other financial qualification exams, financial fundamentals, and financial applications. Law and public administration has two entries: the bar examination and the civil service examination. Reasoning covers deductive, commonsense, and symbolic reasoning via BBH, arithmetic, table question answering and summarization, olympiad mathematics at three levels, sudoku, and several TODO items. Language and instruction following lists idiom understanding, sentiment analysis, textual entailment, text classification, information extraction, reading comprehension, pronoun resolution, poetry matching, Chinese instruction following, and Chinese character glyphs. Agent and tool calling covers TAU and BFCL-V3, and there is a separate coding section with livecodebench and Terminal-Bench-2.0. A final section integrates LMArena and AA scores, which means the project mixes its own measurements with third-party aggregates.

There Is No Install Step in the Material

This is the section where most readers will be disappointed. The README provided here contains no installation instructions, no command line examples, no configuration keys, and no package name. It is a leaderboard document with a table of contents, a changelog, and links. There is no evidence in the supplied material that the repository ships a runnable harness at all. The only operational instructions are social: the README offers free evaluation for private models and directs readers to add a WeChat contact under a heading naming the NoneLinear ReLE benchmark team. So the honest answer to how you get it running is that, based on the repository layout and README, you do not run it. You read the published results, and if you want your own model evaluated you request it through that channel. Anyone claiming a pip install or a docker command for this project is going beyond what the repository states. The homepage at nonelinear.com is listed, and that is where a hosted interface would plausibly live, but the README does not describe one.

The Update Cadence Is the Real Maintenance Cost

The changelog is unusually dense. Between 2025 and 2026 the project moved from v4.3 through v5.11.7, and the recent entries show releases landing every few days: v5.10 on 2026-04-21, v5.9 on 2026-04-18, v5.8.23 on 2026-04-15. Each entry adds models and, notably, deletes them. The v5.11.7 entry removes a long list including gpt-5, gpt-5-mini, gpt-5-nano, several qwen-flash builds, gemini-2.5-flash-lite, DeepSeek-V3.1, claude-haiku-4.5, grok-4-1-fast variants, kimi-k2-0905, claude-sonnet-4.5, and several ERNIE previews. The v5.10.17 entry removes Baichuan4-Turbo, Llama-4-Scout, GLM-4-9B, multiple Qwen3 sizes, o4-mini, DeepSeek-R1-0528, and more. This pruning is a design decision worth naming. It keeps the board readable and current, but it also means a model you evaluated six months ago may no longer appear, and a historical comparison requires archiving the README yourself. If your procurement process references a ReLE ranking, pin the version number and the date in your documentation, because the board is a moving target by construction.

Where ReLE Is the Wrong Tool, and What to Use Instead

ReLE measures models on the maintainers' prompts, with the maintainers' scoring, on the maintainers' schedule. If your question is whether a model handles your own support tickets, your own medical coding guidelines, or your own instruction format, a public leaderboard cannot answer it. The right alternative is a self-hosted harness such as lm-evaluation-harness, which lets you define tasks, run them against a local endpoint, and get per-sample outputs you can inspect. The difference in approach is fundamental: ReLE is a publication, and lm-evaluation-harness is an execution engine. A second alternative is to use the ReLE defect library as a seed corpus and run it through your own harness, which combines ReLE's coverage with your own scoring. That path is only viable if the defect data is machine readable and licensed for reuse, and the README does not state either. A third case: if your product is English-first, the Chinese-language orientation of these domains makes the leaderboard largely irrelevant, and an English-centric suite will serve you better.

Licence, Data Reuse, and What the README Leaves Open

The repository metadata supplied here lists no licence, and the README does not name one. That is a practical obstacle, not a formality. Without a licence you have no stated permission to redistribute the leaderboard data, the defect library, or any derived scores, and the 2 million item corpus is exactly the kind of asset teams want to fold into internal tooling. Before building anything on top of it, check the repository for a LICENSE file and, if one is absent, contact the team through the channel the README provides. The same caution applies to the technical report: it is cited by arXiv identifier, so the paper is the place to look for scoring details, prompt construction, and how the roughly 300 sub-dimensions are weighted into domain scores. The README also integrates LMArena and AA scores into a combined view, and those sources carry their own terms. None of this is legal advice, but the absence of a licence is a concrete blocker worth resolving before the data enters a pipeline.

Editorial conclusion

Adopt ReLE if you need Chinese-language capability comparisons across education, medicine, finance, law, reasoning, instruction following and agent tool calling, and you accept that results arrive on the maintainer's schedule. Do not adopt it if you need a harness you can run against your own prompts, or if your evaluation targets English or non-Chinese-language workloads. Before relying on any number, open the technical report at arxiv.org/abs/2601.17399 and confirm the scoring method, the sample sizes behind each sub-dimension, and the prompt and decoding settings, because the README does not state them.

Official sources

  1. Issues
  2. jeinlee1991/chinese-llm-benchmark on GitHub
  3. Project website
  4. README
  5. Releases
Community notes

Community notes