Static-to-Dynamic-LLMEval is a readme and an image folder, and it says the paper is not the source of truth
The official GitHub repository of the paper "Recent advances in large language model benchmarks against data contamination: From static to dynamic evaluation"
At a glance
- What is it?
- The official repository for a survey of benchmark contamination, twelve authors across eight institutions, organised as a curated bibliography split between static and dynamic evaluation. Its own note says the camera ready cannot be updated in real time, which makes this file the living copy and the paper the stale one.
- Who is it for?
- Use this repository as a map rather than as a tool, because it is a bibliography with a taxonomy attached and nothing executable. Read it for the taxonomy, which is the part the paper cannot keep current, and for the claim that the missing piece is evaluation criteria for dynamic benchmarks rather than another method.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 12 days ago.
- What is it written in?
- GitHub does not report a main language for this repository.
Answers come from the project's GitHub data, last synced on October 10, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The deliverable is a bibliography in a markdown file
The repository contains two entries, a readme and an image folder. There is no source code, no dataset, no evaluation harness, no build configuration and no releases.
That is a normal shape for a survey repository, and this one is explicit about the role. The description says it is the official repository of the paper on benchmark contamination moving from static to dynamic evaluation, and the file itself explains that its purpose is to collect new research as it emerges, to take suggestions on the taxonomy, to catch missed papers and to update preprint entries once an accepted venue is known.
What you get for that is a taxonomy and a list. Every entry follows one format, a paper title, the venue or journal or preprint in italics with the year, and then links to the paper and to any other resources:
Paper Title, <ins>Conference/Journal/Preprint, Year</ins> [[pdf](link)] [[other resources](link)].The italic venue with a year is the field that makes the list useful later, because a preprint becomes a conference paper and a search by title alone loses that. Most entries carry a second link to code, so the repository is a route to implementations as much as to reading.
Two metadata fields are empty: no licence is recorded and no primary language is set, and there is no licence file in the tree. For a bibliography rather than a program that matters less than it would elsewhere, but if you plan to redistribute the list, read the source venues instead of assuming a grant.
The file declares itself the current copy and the paper a lagging one
The most useful sentence in the repository is the note about the paper rather than about contamination. It says that because the authors cannot update the main camera ready in real time, this repository should be referred to for the latest updates, and that the paper may be updated later.
So the artifact with the citation count is the stale one. Camera ready files are frozen for the conference proceedings, and any survey that intends to stay current has to live somewhere else. Stating that openly in the first screen of the file saves a reader from trusting a taxonomy that has already moved on.
The consequence for anyone using the taxonomy is practical. If you cite the paper, cite it for the argument and the structure it had at submission, and check this file for what has been added since. The categories most prone to move are the ones that depend on model capabilities rather than on subject matter.
Contributions are invited through pull requests or by email, with the suggested entry format given so a suggestion is cheap to send, and contributors are told their work will be acknowledged in the paper's acknowledgements. That is a real incentive to participate and also a reason the list will keep growing: the acknowledgement is the only credit a bibliography can offer.
The most recent push to the branch is dated 2026-08-11, a little under two months before this batch was assembled, which fits a file that is maintained between paper milestones rather than on a schedule.
The two halves of the taxonomy are organised by different things
The table of contents splits into static benchmarking and dynamic benchmarking, and the internal structure of each half is not parallel.
The static half is organised by subject area, with eight application groups: math, knowledge, coding, instruction following, reasoning, safety, language and reading comprehension. Below that sits a second axis, four families of mitigation methods, which are canary strings, encryption, label protection and post-hoc detection.
The dynamic half is organised by mechanism instead. It opens with temporal cutoff, then three families of generation: rule-based generation subdivided into template-based, table-based and graph-based; model-based generation subdivided into benchmark rewriting, interactive evaluation and multi-agent evaluation; then hybrid generation; then a network and system category.
So static benchmarks are indexed by what they measure and dynamic ones by how they are produced. That is a defensible choice, since the static side's problems come from its subject coverage while the dynamic side's problems come from its construction, and it is also why a reader looking for work on a particular capability has to look in two different places depending on which half they mean.
The gap the survey names is missing criteria, not missing methods
The argument is worth stating precisely, because it is the reason the survey exists.
The premise is that models trained on web-scraped corpora may have seen benchmark data, which inflates the score and makes the assessment misleading. The claim is that this has become worse with large language models, because they scrape very large amounts of publicly available data, and that the exact training data is difficult to trace, in the survey's words challenging if not impossible, for privacy and commercial reasons.
From there the survey argues that existing work concentrates on static benchmarks with fixed, human-curated datasets, that methods such as encrypting data or detecting contamination after the fact have inherent limitations, and that dynamic benchmarking has emerged as a promising alternative. Then comes the gap: no standardized criteria exist for evaluating dynamic benchmarks.
That framing is the contribution. A survey that catalogued another hundred methods would not change practice. One that says the field lacks the criteria to compare its own methods, and then proposes design principles and evaluation criteria for dynamic benchmarks, is making a claim about infrastructure rather than about artefacts.
The stated contribution list follows the same shape: an analysis of methods that enhance static benchmarks and their inherent limitations, the design principles, an analysis of the limitations of existing dynamic benchmarks, and an overview meant to guide future work.
Four mitigation families, and two of them are named as limited
The mitigation axis under static benchmarking has four families, and they sit at different layers.
A canary string is the oldest and cheapest idea: put a distinctive marker in the test set, then search a model's outputs or training data for it. Encryption and label protection work at the data layer, where labels or the answer key are withheld so a benchmark cannot be reconstructed by scraping. Post-hoc detection comes after the fact and tries to establish from the outside whether contamination happened.
The survey singles out two of them as having inherent limitations, and they are the two most common in practice: encrypting data and post-hoc detection. That is a pointed statement, because those are the two approaches a lab reaches for first when it wants to claim its benchmark is clean.
Detection after the fact is structurally weak, and the survey's own contamination argument explains why. If the training data cannot be traced, no post-hoc method can prove a negative, so detection can only establish a suspicion. Encryption solves the scraping path and does nothing about a benchmark that was already public before it was encrypted, which is the common case for a released dataset.
The dynamic half is the survey's proposed answer, and its taxonomy is long enough to suggest the field is already crowded: three generation families with three subdivisions each, a hybrid category and a network and system category.
The bibliography spans a decade and links to code as often as to papers
The visible entries show the range the survey covers, and each line is dated and linked twice.
The math section lists training verifiers for solving math word problems from 2021, with a link to the arXiv paper and to the code, and the MATH dataset paper from NeurIPS 2021 with its own repository. Knowledge lists TriviaQA from ACL 2017, a paper link into the ACL anthology and a code link, and Natural Questions from TACL 2019 with links to the journal article and to the Google research page, and then a multilingual benchmark paper from ICLR 2021.
Two things are visible from those lines alone. The earliest cited work is from 2017 and the most recent from 2021, so the survey is anchored in the era when contamination became a known problem rather than tracking every development since. And the code links are as prominent as the paper links, which makes the list usable as an implementation index rather than only as further reading.
The entry format also implies curation effort. Venue, publication type and year in one field is the sort of detail that gets filled in by a maintainer as preprints land, and the request for updates about papers that have since been accepted is an admission that this field is one where the venue is the news.
The twelve authors behind the list come from eight institutions across three continents, with a corresponding author listed last, which is the convention when the work is a collaboration rather than a single lab's survey.
Editorial conclusion
Use this repository as a map rather than as a tool, because it is a bibliography with a taxonomy attached and nothing executable. Read it for the taxonomy, which is the part the paper cannot keep current, and for the claim that the missing piece is evaluation criteria for dynamic benchmarks rather than another method. If you want the benchmark itself, this is not where it lives, and the entries are where to look instead: each one links to the paper and to the code that came with it.
Frequently asked questions
Does Static-to-Dynamic-LLMEval contain code or a benchmark?
No. The repository holds a README.md and an image folder. The work is a survey, and its entries link to the papers and to the code released alongside each one.
Why is the repository maintained if the paper is published?
The authors state that they cannot update the camera ready version in real time, so the repository holds the latest updates and the paper may be updated later. The repository is the current copy.
How does the survey organise dynamic benchmarking?
By mechanism rather than by subject: temporal cutoff, rule-based generation with template, table and graph variants, model-based generation with benchmark rewriting, interactive and multi-agent evaluation, hybrid generation, and a network and system category.
What gap does the survey claim to identify?
The lack of standardized criteria for evaluating dynamic benchmarks. It reviews contamination-free strategies, assesses their strengths and limitations, and proposes design principles and evaluation criteria.
Why is training data contamination hard to prove?
Because models scrape very large amounts of publicly available data, and because tracing the exact training data is described as challenging if not impossible for privacy and commercial reasons.
How can I contribute a paper to the Static-to-Dynamic-LLMEval survey?
By pull request or email, using the documented entry format of a paper title, the venue or journal or preprint with the year, and links to the paper and to other resources. Contributors are told they will be acknowledged.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/seekingdream-static-to-dynamic-llmeval)