CSRankings moved its data source after the original started serving a bot challenge
A web app for ranking computer science departments according to their research output in selective venues, and for finding active faculty across a wide range of areas.
At a glance
- What is it?
- A metrics-based ranking of computer science departments that counts publications at selective venues rather than running surveys, with contributions routed through a web form that files an issue and waits for a quarterly processing cycle. Its build notes record a migration to a mirror archive because the primary host began answering a download request with an HTML challenge, and a parser swap that took the working set from eleven gigabytes to fifty megabytes.
- Who is it for?
- CSRankings suits a researcher who wants an affiliation, a scholar identifier and a record of removals handled by someone other than themselves, and who reads a venue-count metric with its stated limitations in mind. It does not suit someone looking for a citable league table, because the project explicitly declines citations for now and describes that as a known gap it intends to close.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The data source moved because the original serves a bot challenge
The most informative comment in this repository is in the build file, and it explains a change that would otherwise look arbitrary.
The download target is not the bibliographic database's own site. It is a different host: a conference archive that distributes the database's releases. The comment says why, and it is specific. Since September 2026 the original hosts sit behind an anti-bot check that answers a command-line download request with an HTML challenge page rather than the archive file. A build that fetched the file with a standard downloader would have silently saved an HTML page and failed much later.
So the project changed its source to keep the pipeline working, and left the reason in the file next to the change.
That is worth noticing for two reasons. First, it dates a real external change to within the last month of this project's life, which is the kind of thing that breaks a build silently. Second, it shows the build has a documented download script of its own rather than a bare fetch line, which is what made the switch a one-line change.
The build also keeps a target for downloading a previous version of the data, and a backup target, so the failure mode of a bad fetch is recoverable.
Eleven gigabytes down to fifty megabytes, in a comment
Two comments in the same build file record the same engineering decision from two directions.
One says the previous approach to filtering the bibliographic data is no longer used. Another says what replaced it: a streaming parser with constant memory, about fifty megabytes, instead of loading the entire document, about eleven gigabytes.
Two hundred and twenty times less memory for the same job, described in two sentences, with the old tool named so a reader knows what it replaced. That is what a well-maintained build file looks like when the person maintaining it has actually run the thing.
The readme says the same thing from the other end, with a different tool named: it recommends a recent Python version and says the filtering uses a streaming parser for efficient memory usage.
The two accounts are consistent and they frame the same requirement. Running this site locally is not a lightweight exercise, and the readme says so with a number: downloading and filtering the data will consume upwards of nineteen gigabytes of memory.
So the streaming parser made the build affordable to run on one machine, and the project still tells you the machine needs a lot of memory. Both facts can be true at once, and it is good that they are both written down.
A no-derivatives licence on a repository that asks you to fork it
The licensing is stated once, at the end of the readme, and it is not the permissive kind.
The project is covered by a Creative Commons attribution, non-commercial, no-derivatives licence at version four international. The build file's header points at a copying file in the root for licence information, and its own copyright line reads 2017 to 2020.
That combination deserves a second look, because a no-derivatives clause restricts redistribution of adapted material, and the readme spends a section teaching people how to fork the repository, clone it shallowly, edit the data files, and open a pull request. Forking is derivative use by another reading.
The data has a second licence layered on top. The readme states that the bibliographic information the site uses is made available under a separate attribution licence from the database it comes from. So there are two terms in play: the site's own content under a restrictive Creative Commons variant, and the upstream data under an attribution-only licence.
Nothing here is unusual for an academic data project. The point worth making is that anyone planning to republish the rankings, build a derivative dataset, or combine them with another source needs to read both licences rather than assuming the repository's terms cover the data.
Contributions go through a form, on a quarterly cycle
The contribution instructions are unusual, and the reason is in the first line of them.
It says essentially all contributions should be made through a self-service submission form, and then lists what that avoids: no clone, no editing the comma-separated data files, no worrying about the file format. The form validates the entry against four things, a bibliographic database name, a homepage, a scholar identifier, and an identifier from the research community, and then files an issue that is processed automatically into a pull request.
So the human contribution path is a web form, and the code path that most contributors would expect, editing a file and opening a pull request, is explicitly reserved for maintainers and unusual cases.
The cycle is the other fact. Submissions are processed on a quarterly basis, and the readme says outright that a change may take up to three months to appear.
There is also a routing rule for a new institution: open an issue from a specific template first, then submit its faculty through the form. And one file is off limits entirely, with a note saying it is generated automatically.
For a dataset with thousands of rows and a volunteer maintainer, a validated form plus an automatic pull request is a defensible design. The three month wait is the cost of it.
Removals are archived, and the company is recorded
The removal path is the part of the form worth reading twice, because it handles the cases that usually get argued about.
The form covers six situations: someone who retired, someone who became an emeritus professor, someone who moved to industry, someone who is deceased, someone who is no longer on a tenure track, and someone who left academia entirely.
Two details in the industry case. The record stays, and the company is written down alongside it. So the dataset keeps the faculty member and records where they went, rather than deleting them.
And the readme is explicit that entries are moved into an archive directory rather than deleted. Removal is a move, not a deletion.
That matters for a rankings site, because the metric is a count of publications at selected venues attributed to current faculty. If a record were deleted outright, a past ranking would no longer be reproducible from the current data. Keeping the record with a status, and keeping removed records in an archive, is what lets a ranking from three years ago be checked against the data that produced it.
The same form handles movement between institutions, updates to a homepage or scholar identifier, and a disambiguation suffix for faculty whose names collide in the bibliographic database, written as four digits appended to the name.
Twenty-six lettered data files, and one that is generated
The root of this repository is mostly data, and the way it is split tells you how it is maintained.
There are twenty-six comma-separated files, one per letter of the alphabet, plus the aggregate file. The aggregate is the one the readme tells you not to edit, because it is generated. The lettered files are the ones maintainers work in, and direct edits to them by pull request are reserved for maintainers and unusual cases.
So there are two layers: a generated aggregate for the site to consume, and hand-maintained shards that people can actually review. Twenty-six shards is a lot of files to review, which is presumably the other reason the form exists.
Around them sit the inputs and the by-products. A country lookup, a file of distinguished fellows, an aliases file that the build regenerates from the downloaded data, and the schema file for the bibliographic format. The compressed raw data dump is committed too, along with its image and a small related file that looks like an advertising verification record.
Then the things a normal repository would not have at the root: a citation file, a domain file for the hosting, four planning and validation documents, an analysis script that renders to both a document and an image, and two pages of prose advice and frequently asked questions.
You also need to install the following dependencies: ```bash apt-get install libxml2-utils npm npm install -g typescript google-closure-compiler python3 -m pip install -r requirements.txt ``` Two Python variables point at the same interpreter
The build file defines two Python variables and assigns the same value to both, which is a small artefact of how the project used to work.
One is the interpreter to use. The other is named for an alternative Python implementation and is assigned the same value. Both carry a trailing comment showing what they used to be: an older Python version for the first, and that other implementation for the second.
The second variable is not referenced by any target in the visible portion of the file. So there is a variable that exists only to hold an interpreter nobody runs, and its name suggests a capability the project no longer exercises.
The comment above the data source explains why the second variable went away in practice. The original plan was a database engine for filtering the bibliographic data; that was replaced by a streaming parser, and with it went any reason to run under a second implementation.
That is a normal fossil. It is worth mentioning only because it is the kind of thing that trips up a contributor trying to run the build with the interpreter the project used to prefer, and the comment does not tell them it is no longer used.
Editorial conclusion
CSRankings suits a researcher who wants an affiliation, a scholar identifier and a record of removals handled by someone other than themselves, and who reads a venue-count metric with its stated limitations in mind. It does not suit someone looking for a citable league table, because the project explicitly declines citations for now and describes that as a known gap it intends to close. Before you rely on a rank, check three things: how recent the affiliation data is, given submissions are processed quarterly and a change can take up to three months to appear; whether your area is covered by the venue lists the project curates; and what the no-derivatives licence on the repository means for anything you build on top of the underlying data.
Frequently asked questions
what is cs rankings
An entirely metrics-based ranking of computer science departments and faculty. It counts the number of publications by faculty that appeared at the most selective conferences in each area of computer science, and it positions itself against survey-based rankings by saying it derives its numbers from venue lists rather than from questionnaires. It is also used to find active faculty across a wide range of areas.
is csrankings reliable
The project's own answer is methodological rather than a claim of accuracy. It says the approach is intended to be difficult to game because publishing in those venues is generally hard, contrasts this with citation-based metrics as repeatedly shown to be easy to manipulate, and then concedes that incorporating citations in some form is a long-term goal. It points at a frequently asked questions page on the site for more detail.
How accurate are csrankings?
The repository publishes no accuracy assessment. What it does document are the mechanics: the submission form validates a bibliographic name, a homepage, a scholar identifier and a research identifier, and files an issue that is processed into a pull request; submissions are processed quarterly so a change may take up to three months to appear; and removals move records into an archive directory rather than deleting them.
csrankings vs us news
The readme frames the difference as methodology rather than as a comparison of results. The survey-based approach is described as exclusively survey driven, while this project counts publications by faculty at the most selective conferences in each area. No side-by-side ranking comparison is offered anywhere in the repository, and the difference in results between the two is not something the project claims.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/emeryberger-csrankings)