# anonymous_github: hiding an author's identity from a code review

> Double-anonymous peer review asks you to anonymise the code behind a paper, not just the paper. This proxy does that mechanically, and the interesting parts are the credential key rotation and the hidden replica it can run.

**tdurieux/anonymous_github** — Anonymous Github is a proxy server to support anonymous browsing of Github repositories for open-science code and data.

- Repository: https://github.com/tdurieux/anonymous_github
- Website: https://anonymous.4open.science/
- Stars: 2,251 · Forks: 97
- Language: JavaScript
- License: GPL-3.0
- Published: 2026-10-07 · Updated: 2026-10-07 · Language: en
- Canonical page: https://hysenlabs.com/projects/tdurieux-anonymous-github

## The problem is mechanical, which is why a tool can fix it

The stated motivation is precise. Double-anonymous review asks you to anonymize the artifact behind your paper, the code and data, exactly like the paper itself. Doing that by hand means going after the owner, the organization, the repository name, logins, and every mention buried in the files. The README's word for that process is tedious and easy to get wrong, and both halves of that are true: it is repetitive, and the failure mode is a single missed string that identifies the author.
What gets replaced is listed explicitly, and the list is narrower than you might expect. The GitHub owner, organization, and repository name. File and directory names. And the file contents of every extension, covering Markdown, text, and source code, with a configurable list of terms to replace. That last one is the design centre of the tool. Everything else is a rename; the content replacement is the part that requires understanding what identifies a person, such as a name in a copyright header, a university path in an import, or an acknowledgement section.
Two usage paths are offered and they differ in where the anonymized repository lives. The public instance at anonymous.4open.science takes a pasted repository URL plus the terms to hide and returns a shareable link, so nothing is uploaded by you. The CLI anonymizes a repository locally and produces an anonymized zip, which is the path to take if the artifact is not on GitHub, if the repository is private, or if you would rather not hand it to a third party at all:
```bash
npm install -g @tdurieux/anonymous_github
anonymous_github
```
Both routes end in the same place, a link or a zip that a reviewer can open without learning who wrote it.

## The anonymization boundary is drawn deliberately

A section on scope of anonymization states the boundary plainly: in double-anonymous review, the boundary is the paper plus its online appendix and only that. Googling part of the paper or appendix to reveal authorship is considered a deliberate attempt to break anonymity, with a link to an explanation of the reasoning.
This is the kind of scope statement most tools omit and every reviewer needs, because it determines whether a tool is doing its job. If you upload a repository that contains material outside the appendix, such as a companion website or a separate analysis repository, then a searcher who knows to look there will find the author regardless of how thoroughly the submitted repository was scrubbed. The tool cannot fix that, and pretending otherwise would be dishonest.
It also sets an expectation about what the reviewer does next. A well-anonymised artifact still carries the fingerprints that only the authors know to look for, and the README is clear that actively searching for authorship is treated as misconduct rather than as diligent reviewing. That is a policy statement rather than a technical feature, but it belongs in the same document because it tells an author what protection they are actually getting.
Two related tools are listed for the adjacent problems. gitmask is for contributing anonymously to a GitHub repository, which is the mirror image of this problem: hiding the contributor rather than the repository owner. blind-reviews is a browser add-on that hides identifying information when reviewing a pull request. Together the three cover submission, contribution, and review, which is a sensible way for one project to delimit itself.

## Credential encryption with dated keys and a migration path

The self-hosting instructions contain the most production-minded configuration in the README, and it is easy to skim past. The .env file includes a CREDENTIAL_KEYS entry that is a JSON object keyed by a date string, with a base64-encoded 32-byte random key as the value, and a separate CREDENTIAL_ACTIVE_KEY_ID naming which key encrypts new credentials:
```env
GITHUB_TOKEN=<GITHUB_TOKEN>
CLIENT_ID=<CLIENT_ID>
CLIENT_SECRET=<CLIENT_SECRET>
CREDENTIAL_KEYS='{"2026-09":"<base64-encoded 32-byte random key>"}'
CREDENTIAL_ACTIVE_KEY_ID=2026-09
CREDENTIAL_LEGACY_READS=false
PORT=5000
```
The design is envelope encryption with explicit key versioning. A key named 2026-09 writes; older named keys still decrypt what they encrypted, so rotating does not orphan existing credentials. The README instructs you to generate a key with openssl rand -base64 32, and the warning is that existing installations must follow the credential migration guide before starting this release, with a matching migrate:credentials npm script in the source.
CREDENTIAL_LEGACY_READS is the switch that makes the rotation a real decision rather than a formality. Setting it false means the server will not fall back to reading credentials in the old unencrypted form. Leaving it true would preserve compatibility at the cost of keeping the insecure path alive. Making that a named environment variable rather than a code change is the right call, because it means an operator can see the decision in their own configuration file and can grep their fleet for installations that still have it enabled.
The rest of the authentication configuration follows the standard GitHub OAuth App shape: a client id and secret from an OAuth App, a callback URL that must be https://<host>/github/auth, and a GITHUB_TOKEN with the repo scope created at the tokens settings page. There is also a GitHub App setup guide for read-only access to selected private repositories alongside OAuth, which covers the case where you want to mirror a private artifact into an anonymised public copy without granting blanket repository access.

## A streamer pool in front of MongoDB and Redis

The compose file is where the architecture becomes clear, and it is more than a single service. There is the main anonymous_github service with a 3G memory limit, a streamer service defined with four replicas behind DNS round-robin and a 768M limit, and then Redis and MongoDB. The main service waits for all three to report healthy before starting, using condition: service_healthy on each dependency, so a cold start cannot race.
The streamer exists because the expensive work is downloading and rewriting repositories, which is CPU and memory bound and variable in duration, while serving a page for a reviewer is neither. The compose entrypoint for the streamer runs the built streamer entrypoint with a 512MB old-space cap, and it points at the main service through STREAMER_ENTRYPOINT. That split is what lets the deployment scale the slow work independently of the read path.
The runtime container is deliberately leaner than the build container, which is a sign of a deployment that has been tuned rather than assembled. The image builds on node:22-slim for the dependency install, the TypeScript and Gulp asset build, and a production-only dependency install, then switches to node:22-alpine for the runtime stage, setting NODE_ENV to production and PORT to 5000. Each stage copies only what it needs from the previous one, and the build stage copies just package.json, package-lock.json, tsconfig.json, gulpfile.js, public, and src so that a source-only change does not invalidate the dependency cache.
Both the npm and the Docker caching use buildkit cache mounts, which is what makes the multi-stage build fast on repeat deploys. The healthcheck is a node healthcheck.js script run every five seconds with a five second timeout and three retries, and the compose file applies the same check to the streamer. Given that the health state gates the depends_on conditions, that script is load-bearing rather than decorative.

## A hidden delayed replica for the MongoDB data

The README mentions an optional remote, hidden, delayed MongoDB replica and backup source, with a separate replication guide in the docs directory. The tree contains three compose files beyond the default one: docker-compose.yml, docker-compose.github-app.yml, and a replica pair split into docker-compose.replica-primary.yml and docker-compose.replica-secondary.yml.
This is the part of the design that deserves the most thought, because it addresses a trust problem rather than a technical one. A tool whose entire purpose is protecting an author's identity is operated by somebody, and that operator can read the MongoDB, which holds the mapping from anonymized repository back to the real one along with the cached originals. A remote replica that is hidden from the operator and delayed means that even a compromised or curious operator cannot immediately read submissions, and the delay bounds how fast a leak can be exfiltrated.
Delayed replication is a well-established technique from journalism and from whistleblowing infrastructure, and applying it to peer review submissions is a direct transfer. The cost is that recovery from the replica lags behind, so if the primary is lost you may lose recent submissions, and if a legal request arrives the operator can honestly say the data is not currently present while it still arrives eventually. Neither property is free, and the README points you at the guide rather than assuming you will infer the trade.
The README also flags that the original and anonymized versions are cached on the server in both operating modes, whether downloading and rewriting files or proxying on the fly, so that even large repositories stay responsive. Read together with the hidden replica, the picture is consistent: caching is treated as necessary for usability, and the mitigation for the resulting privacy exposure is architectural, keeping the cache somewhere the operator cannot read rather than pretending the cache is not there.

## Vue 3 frontend, TypeScript backend, and a JavaScript label

The repository is labelled JavaScript and is in practice a TypeScript backend with a Vue 3 frontend. The evidence is in the build: tsconfig.json at the root, a build script that removes the build directory and runs tsc and then gulp, dev and start scripts that use ts-node/register to run ./src/server/index.ts directly, and two maintenance scripts, migrate:credentials and recover:owners, that also run TypeScript through ts-node. The repository language is detected from what dominates the file count, and the plain JavaScript in public/, the Gulp file, and the eslint config outweighs the src/ tree.
The frontend section explains how the two halves meet. The UI uses Vue 3 and Vue Router against the existing Express API. Page setup functions live in public/script/app.js and admin.js, the Vue templates are in public/partials/, and Gulp compiles the templates and bundles the app with esbuild before updating the asset manifest Express serves. The workflow instructions are correspondingly direct: run npm run build:ui after changing a template or frontend script, and npm run dev:ui to serve the built UI locally with API requests proxied.
Server-rendered templates compiled by Gulp rather than a single-page bundler is a deliberate and slightly old-fashioned choice, and it has one practical advantage: the public/ directory holds what the browser receives, so the served HTML is inspectable without a build step running first. It also means the Vue components are templates rather than single-file components, which some teams prefer and others find limiting.
The test story is more interesting than the framework choice. Alongside mocha with the spec reporter and c8 for coverage scoped to src/core, src/server/routes, and src/config, there are three frontend test scripts named frontend-regressions, dashboard-ui, and vue-ui, run through gulp first. Naming a test file frontend-regressions is a statement of intent, and it suggests the team hit UI breakage and institutionalised the check. There is also knip for finding unused dependencies and exports, healthcheck.js at the root for the container check, and a test-mermaid.html paired with a test-mermaid.md sitting at the repository root, which are the kind of fixture pair that documents a rendering path for diagram content. This project anonymises Markdown files, and a rendering test for Mermaid suggests someone needed to know what happened to diagrams inside documents.

## Conclusion

This is a well-engineered answer to a problem that is genuinely tedious rather than merely inconvenient, and the parts that matter most are the unglamorous ones. Credential keys are versioned by date, encrypted at rest, and rotated through an explicit active key id, with a migration script and a documented guide for existing installations. The deployment splits into a stateless streamer pool and a main service backed by MongoDB and Redis, and it offers an optional remote replica that is hidden and delayed, which is the correct design for a tool that handles unpublished manuscripts. Two limits to weigh. The README states the anonymization boundary is the paper plus its online appendix and only that, so this will not strip identity from a linked dataset you control elsewhere. And the service caches both the original and the anonymised repository on its server, which means whoever runs the instance holds the unanonymised source of everything submitted to it.

## FAQ

### What does anonymous_github anonymize?

It replaces the GitHub owner, organization, and repository name, renames file and directory names, and rewrites the contents of text-based files including Markdown and source code using a configurable list of terms. Binary files are not rewritten.

### Do I need to upload my repository to run it?

Not necessarily. There is a CLI that anonymizes a repository locally and produces an anonymized zip, so nothing has to leave your machine. The public instance is the alternative, where you paste a repository URL and receive a shareable link.

### How are stored GitHub credentials protected?

They are encrypted with keys stored in CREDENTIAL_KEYS as a JSON object keyed by a date, with CREDENTIAL_ACTIVE_KEY_ID naming the key used for writes so older keys can still decrypt what they encrypted. CREDENTIAL_LEGACY_READS controls whether the server will fall back to reading credentials in the old unencrypted form.

### Can I self-host anonymous_github?

Yes. You clone the repository, run npm i, create a .env file with a GitHub token and OAuth app credentials, then start it with docker-compose up -d. The README recommends putting the service behind nginx for HTTPS and changing the exposed port in the compose file.

### What is the streamer service for in the Docker setup?

It runs the download and anonymization work separately from serving pages, with four replicas in the default compose file and its own memory limit. The main service waits for the streamer, MongoDB, and Redis to report healthy before starting.

## Sources

- [Issues](https://github.com/tdurieux/anonymous_github/issues)
- [License: GPL-3.0](https://github.com/tdurieux/anonymous_github/blob/main/LICENSE)
- [Project website](https://anonymous.4open.science/)
- [README](https://github.com/tdurieux/anonymous_github/blob/main/README.md)
- [tdurieux/anonymous_github on GitHub](https://github.com/tdurieux/anonymous_github)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/tdurieux-anonymous-github
