# LLM-eval-survey is a paper list, not a tool, and its news log stopped in 2023

> The official GitHub page for a survey on large language model evaluation, holding sixteen authors' worth of categorised reading rather than code. The repository exists because the arXiv paper cannot be updated, and its own update section has not moved since December 2023.

**MLGroupJLU/LLM-eval-survey** — The official GitHub page for the survey paper "A Survey on Evaluation of Large Language Models".

- Repository: https://github.com/MLGroupJLU/LLM-eval-survey
- Website: https://arxiv.org/abs/2307.03109
- Stars: 1,612 · Forks: 105
- Language: Unknown
- License: not declared
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/mlgroupjlu-llm-eval-survey

## The update log exists because the paper cannot change, and it stopped in 2023

The stated reason this repository exists at all is that the survey cannot be revised in real time.

> As we cannot update the arXiv paper in real time, please refer to this repo for the latest updates and the paper may be updated later.

By that logic the news section is the most important part of the file, and it is the smallest. It has exactly two entries. The first records the first arXiv version on 07/07/2023, and the second records the second version on 12/07/2023 together with a Chinese blog post linked from Zhihu. Nothing has been added to it since. The repository itself has not been idle, since the last recorded commit on the main branch is dated 2026-09-13, and the project publishes no GitHub releases, so there is no tag history to compare against the list either. A reader who trusts the heading rather than the entries will assume the list tracks the field, and the heading says it is the place to look for what is newest.

## The same paper appears under several headings, so counting entries overcounts

The list is organised as a set of numbered subsections under each branch of the taxonomy, and each subsection restarts its numbering at one. Papers are not exclusive to a subsection. The Percy Liang paper on evaluating language models holistically is listed under sentiment analysis and again under text classification. The Laskar study of ChatGPT across benchmark datasets appears under sentiment analysis, under natural language inference, and again under reasoning. The Chengwei Qin paper on ChatGPT as a general purpose natural language processing task solver is listed twice, under sentiment analysis and under natural language inference. This is defensible as a reading list, since a paper that spans two task types is genuinely relevant to both, but it means the total number of numbered items in the file is not the number of distinct papers. Anything that counts the list programmatically, to claim coverage of a subfield or to build a reading queue, has to deduplicate by title or by link first.

## The table of contents link for Where to evaluate is missing its hash

The navigation block is built by hand inside a details element, as an ordered list of anchors. Every entry points at a fragment on the page except one.

```html
<li><a href="#what-to-evaluate">What to evaluate</a></li>
<li><a href="where-to-evaluate">Where to evaluate</a></li>
```

The first is written with a leading hash and the second is not, so the Where to evaluate entry resolves as a relative path instead of an in page anchor. Given that the contributing, citation and acknowledgements entries in the same list all carry the hash, this looks like a slip rather than a different link target. The structure of the list is otherwise worth reading as a statement of the survey's shape: what to evaluate branches into seven areas, natural language processing, robustness and ethics and biases and trustworthiness, social science, natural science and engineering, medical applications, agent applications and other applications, and then where to evaluate sits beside it as the second half of the taxonomy.

## The two related projects are separate efforts, and one link label is misspelled

Two items sit under related projects. The first is PromptBench, described as robustness evaluation of large language models and hosted in the Microsoft organisation. The second is an evaluation site at llm-eval.github.io, and its label is written with the letters transposed, reading Evlauation of large language models. Both are pointers to work maintained elsewhere; neither is vendored into this repository, and neither is described here beyond a line of text. That matters for anyone who lands on the page looking for a harness: the top level tree holds a README file and an image directory, so there is no package manifest, no requirements file, no notebook and no evaluation code to run. The page is a bibliography with a taxonomy attached, and the two links are the only executable-looking things in it.

## Sixteen authors across eight institutions, marked for first and corresponding roles

The author block uses numeric superscripts for affiliations and two letters for roles. Yupeng Chang and Xu Wang carry the asterisk for co-first authors, and Jindong Wang carries the hash for co-corresponding authors. The footnote under the block spells out that asterisk and hash, which is the only place the convention is defined. Eight institutions are numbered: Jilin University, Microsoft Research, the Institute of Automation at the Chinese Academy of Sciences, Carnegie Mellon University, Westlake University, Peking University, the University of Illinois, and the Hong Kong University of Science and Technology. Three of the sixteen authors are marked as affiliated with Microsoft Research and four with Westlake University, so the centre of gravity is spread across a Chinese university group with two industry and US academic anchors. That level of attribution is unusual for a reading list and is the main reason the page works as a citation target rather than a link dump.

## The taxonomy is keyed to task type, so a medical paper sits under reasoning

The seven branches under what to evaluate are organised by what is being measured rather than by which community does the measuring, and the paper list shows where the seams fall. Medical applications has its own branch, yet the paper on whether large language models can reason about medical questions is filed under reasoning, next to commonsense and chain of thought work, not under the medical branch. Robustness, ethics, biases and trustworthiness is one branch holding four topics that a lab would usually split, while social science and natural science and engineering each get a branch of their own. Inside the natural language processing branch, the visible entries run from sentiment analysis through text classification, natural language inference and others, then into reasoning. That is a judgement about where a reader would look for the paper, which is a reasonable choice for a reading list and a debatable one for anyone counting coverage per discipline.

## No license file and no recorded license, for a list people want to reuse

The repository declares no license. The metadata records the license field as unknown, and the top level tree contains a README file and an image directory with no license document among them. For a bibliography that is an unusual gap, because the one thing a reader of a curated paper list usually wants to do is take it somewhere else, as a reading group syllabus, a course reading list or a starting point for a new survey. Without a license file, none of those uses is granted by the repository, and the paper it accompanies carries its own separate terms on arXiv. Contributors are asked to send pull requests or issues and are told their contributions will be acknowledged in the acknowledgements section, so the file is designed to be edited by others, but the permission side of that arrangement is left unstated. Anyone planning to mirror the list should raise that question with the maintainers rather than assume the survey's own licence covers it.

## Conclusion

Use this repository when you want the survey's taxonomy rather than a runnable harness, because the whole artefact is a single Markdown file and the tree contains nothing else to install or call. Two things to check before relying on it as a bibliography. The section that exists to carry what the paper cannot, a news and updates log, has two entries and both are from 2023, so treat the list as a snapshot from that period rather than a current map of the field. And there is no license file at the root and no license recorded in the repository metadata, so reusing the curated list as a dataset or a starting bibliography is a decision you have to make yourself rather than something the project grants.

## FAQ

### What is the LLM-eval-survey repository for?

It is the official GitHub page for the survey paper A Survey on Evaluation of Large Language Models, arXiv 2307.03109. The papers are organised by what to evaluate and by where to evaluate.

### How current is the LLM evaluation survey list?

Its news section has two entries, the first arXiv version on 07/07/2023 and the second on 12/07/2023. The last recorded commit on the main branch is dated 2026-09-13, and no GitHub releases exist.

### Does the LLM-eval-survey repository contain any code or data files?

No. The top level tree holds only README.md and an image directory, so there is no package manifest, requirements file or notebook. The repository is a categorised paper list.

### Which related projects does the LLM-eval-survey page link to?

Two. PromptBench, described as robustness evaluation of large language models and hosted in the Microsoft organisation, and an evaluation site at llm-eval.github.io whose label is misspelled in the page.

## Sources

- [Issues](https://github.com/MLGroupJLU/LLM-eval-survey/issues)
- [MLGroupJLU/LLM-eval-survey on GitHub](https://github.com/MLGroupJLU/LLM-eval-survey)
- [Project website](https://arxiv.org/abs/2307.03109)
- [README](https://github.com/MLGroupJLU/LLM-eval-survey/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/mlgroupjlu-llm-eval-survey
