# A readme that is regenerated daily from a spreadsheet, with liveness shown as emoji

> This list of onion services for mainstream sites is not maintained by hand: the readme is generated, the liveness of every address is probed on a schedule, and each entry carries a two-week history of check results as tooltip text. The design decision worth understanding is that the liveness data is in the document at all, which makes the list's own reliability auditable rather than asserted.

**alecmuffett/real-world-onion-sites** — This is a list of substantial, commercial-or-social-good mainstream websites which provide onion services.

- Repository: https://github.com/alecmuffett/real-world-onion-sites
- Stars: 2,241 · Forks: 193
- Language: Python
- License: not declared
- Published: 2026-09-30 · Updated: 2026-09-30 · Language: en
- Canonical page: https://hysenlabs.com/projects/alecmuffett-real-world-onion-sites

## Inclusion policy first, because the list is an argument about what counts

The readme opens with its own rules, and they are worth reading before the entries. No sites with an onion-only presence, which excludes anything that exists on Tor and nowhere else, on the grounds that a mirror of a site you can already reach is a different thing from a service. No products or technologies with fewer than an arbitrary ten thousand users, and the word arbitrary is in the original, so the editor is telling you the threshold is a judgement rather than a measurement. No nudity, exploitation, drugs, copyright infringement or sketchy-content sites. And then the sentence that decides everything else: the editor reserves all rights to annotate or drop any or all entries as deemed fit. That is an editorial policy, not a technical one, and it means the list is a curated document with an opinion. The practical consequence for a user is that absence means nothing. If a service you want is not in the list, that is not a statement about the service, and if an entry is annotated, the annotation is the editor's judgement rather than a measurement. The stated licence is Creative Commons share-alike, attributed to a named editor, and the readme is explicit that it is not software: the primary language field records Python, but what you would use is the generated document.

## Every entry carries fourteen days of liveness checks

This is the feature that makes the list worth reading rather than a static mirror. Each entry has a check line, and the line is a sequence of emoji spans whose tooltip text holds the detail: an attempt count, a numeric code, an exit code and a timestamp. A single entry in the readme shows a run of results across fourteen consecutive days, mostly the same code and exit status with occasional variations, and each variation is a different emoji. The visible difference is a single character; the diagnostic content is in the hover text. So the readme encodes, for every address, a recent history of what a prober saw. Three consequences follow. First, you can distinguish a site that has been down for a week from one that was down once, which a green tick cannot tell you. Second, the codes are exposed rather than summarised, so someone who knows what the numbers mean can read the history without running anything. And third, the list is falsifiable: if the emoji disagree with your own experience, you have evidence. The cost is that the visible signal is coarse and the tooltips are the real data, which is an odd interface for a document people read at a glance. It is also why the timestamps matter, since the newest entry shown is from the day before the last commit. A list that shows liveness without showing how liveness was measured is just an opinion with a tick, and this one shows the raw numbers.

## Transport is a first-class field, because HTTP and HTTPS onion addresses differ

Each entry begins with a transport line stating whether the service is reached over HTTPS or plain HTTP, marked with distinct symbols. This is a small design decision with a real consequence. An onion address is a hostname resolved inside the Tor network, and a site that also serves plain HTTP over that address is not providing the end-to-end encryption you get from HTTPS, which means the exit and the path are the only protection rather than the path and the destination. Listing the transport lets a reader see, before clicking, whether the connection is encrypted, and it lets a reader judge whether that matters for what they are doing. The other per-entry fields follow a consistent pattern. There is a link, which is a clickable form of the address. There is a plain form in code formatting, so you can copy it accurately. And there is a proof link, which is a page on the clear web belonging to the site owner, so a reader can confirm the operator has published the mapping between their normal address and this one. That proof field is the anti-impersonation mechanism: an onion address alone proves nothing about who runs it, so the list defers to a clear-web statement from the operator. Any tool that consumes this list should treat the proof link as the trust anchor and the liveness as a health check, not the other way round.

## Sections are shaped around who runs the site, not what it does

The index at the top of the readme is the clearest statement of the editorial approach, and it is unusual. The categories are blogs, civil society and community, education, government, news, three separate news subsections for specific broadcasters, search engines, social networks, tech and software, web and internet, and SecureDrop. Read as categories of subject matter, that is a fairly conventional directory. Read as a list, it says something: government services and news organisations have onion addresses, three international broadcasters are broken out into their own subsections, and there is a dedicated section for secure submission systems. That is a list shaped around the set of organisations for which anonymity and circumvention are legitimate, non-criminal concerns, which is a more interesting editorial line than a generic link dump. The three news subsections are the tell. Splitting one organisation out of a news section implies the entries are being tracked individually, which means the editor expects them to move and wants readers to notice when they do. The SecureDrop section is different again and is handled by a separate mechanism entirely, which is the subject of the next section.

## SecureDrop is imported, and the readme tells you where to change it

One entry in the update instructions is a boundary rather than a procedure. The readme says that for SecureDrop, all entries are taken automatically from an external directory API and must be amended on that site, not on this one. So a section of the list is not editorial at all; it is a sync of somebody else's data, and changes flow one way. That is a good decision, because a directory of secure submission systems is a document with real consequences and a single authoritative source is better than a maintained copy, but it has an implication for anyone reading the list. A SecureDrop entry here is a mirror, and the authoritative version, including anything about a site's status, lives elsewhere. The rest of the update instructions are equally direct. The readme is auto-generated from a spreadsheet, which is the central fact about this project. Suggestions go through an issue, and the readme says in bold not to submit changes and not to submit pull requests. So there is exactly one write path, and it is not the normal one for a repository. A reader who wants a site added files an issue and waits for an editor; a reader who wants a site removed has the same route. For a fork, the practical consequence is that you inherit the generated document and the scripts but lose the spreadsheet, and the list will stop updating unless you reconstruct the source of truth.

## The scripts show how the daily update actually runs

The top-level listing and the makefile together describe the pipeline, and it is worth reading because it explains why the document looks the way it does. The update target is the whole workflow in four lines:

```bash
git pull
./wrapper.sh
./get-ct-log.sh
git push
```
 There is a wrapper script that is the main entry point, a shell script for dumping a site's contents, a manual check script, two scripts for fetching a certificate transparency log and a fresh copy of the spreadsheet, a Python script for the certificate transparency log, another Python script for pulling in the SecureDrop directory into a comma-separated file, and yet another for a database. There is also a changelog file in two forms, a makefile, and the spreadsheet itself committed as a comma-separated file. The makefile is where the design decision shows. The default target just echoes a question, because the real workflow is one named run target, and that target pulls changes, runs the wrapper, fetches the log, stages everything, commits with a message containing the current date, and pushes. There is a separate target for rotating logs, one for opening a local database, and one that destroys the database and the master spreadsheet, which is the reset button for a botched run. So this is a single-editor, single-machine, commit-and-push pipeline rather than a continuous integration job, and that is consistent with everything else in the repository. If you are wondering whether the list is reproducible, the answer is that it is reproducible in the sense that a person can run it again, not in the sense that anyone can.

## Conclusion

Use this list if you need to reach a mainstream service over Tor and want to know whether the address was working this morning rather than whether it worked when somebody added it, because the two-week check history attached to every entry is a liveness record rather than a claim. Do not automate against it, since the readme is regenerated from a spreadsheet, pull requests are explicitly not accepted, and the update path is a single editor with a script that commits and pushes, which means a fork cannot easily stay current. Three things to understand. That the inclusion policy is editorial and narrow, excluding any site with an onion-only presence, any product or technology site under an arbitrary ten thousand users, and a stated list of content categories, with the editor reserving the right to annotate or remove entries. That the SecureDrop section is not maintained here at all, since those entries are taken from an external directory and must be changed at the source. And that the liveness indicators reflect a probe from one vantage point, so a failure means that probe failed, which is not the same as the site being down. The readme states the content licence as CC-BY-SA, the last push was on 2026-09-28, and no releases are published.

## FAQ

### What is in the real-world-onion-sites list?

Onion addresses for substantial mainstream websites that are commercial or socially beneficial, grouped into blogs, civil society and community, education, government, news, specific news broadcasters, search engines, social networks, tech and software, web and internet, and SecureDrop. The readme says no site with an onion-only presence is included, and the editor reserves the right to annotate or remove entries.

### How do I know whether an onion address in the list is working?

Every entry has a check line with one symbol per day for the last two weeks, and hovering reveals the attempt count, numeric code, exit code and timestamp behind each symbol. That makes the liveness history auditable rather than a single pass or fail indicator.

### Why does the list say not to send pull requests?

Because the readme is auto-generated from a spreadsheet, so hand edits would be overwritten. The readme asks for an issue instead, and the update path is a make target that runs the scripts and then commits and pushes the regenerated file.

### How are SecureDrop entries maintained?

Not in this repository. The readme says all SecureDrop entries are taken automatically from the external directory API and must be amended on that site rather than here, so that section is a one-way sync of somebody else's data.

### What licence is the real-world-onion-sites list under?

Creative Commons share-alike, as stated in the readme, with the editor named. No standard software licence is recorded, and the repository publishes no GitHub releases, with the last push on the master branch being 2026-09-28.

## Sources

- [alecmuffett/real-world-onion-sites on GitHub](https://github.com/alecmuffett/real-world-onion-sites)
- [Issues](https://github.com/alecmuffett/real-world-onion-sites/issues)
- [README](https://github.com/alecmuffett/real-world-onion-sites/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/alecmuffett-real-world-onion-sites
