Open-source project
binhnguyennus/awesome-scalability avatar
binhnguyennus/awesome-scalability

awesome-scalability is a link list indexed by symptom, not by topic

The Patterns of Scalable, Reliable, and Performant Large-Scale Systems

74,011 stars7,126 forksUnknownMIT

At a glance

What is it?
This is a curated reading list for scalable systems, organised into eleven sections and entered by the reader's situation rather than by subject: a system that is slow, a system that is down, a system design interview, a growing team. Every entry is a link to somebody else's article, none carries a date or an annotation, and the project asks readers to file a pull request when a link dies.
Who is it for?
This list suits an engineer who already has a symptom and wants reputable places to read about it, and the symptom-based entry points are the part worth borrowing whatever else you decide. It does not suit someone building a curriculum, because entries are unannotated titles with no dates, and it does not suit anyone who has to know how fresh a link is, because nothing in the repository records when an entry was last checked.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 8 months ago.
What is it written in?
GitHub does not report a main language for this repository.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

Editorial analysis

You enter the list by symptom, not by subject

The organising idea is unusual and it is the most useful thing in the repository. Four entry points sit at the top, each aimed at a state the reader is in. If the system goes slow, the text asks you to work out whether you have a scalability problem, fast for one user and slow under load, or a performance problem, slow for one user, and sends you to the design principles, the scalability and performance sections, and the intelligence section. If the system goes down, it points at availability and stability. If there is a system design interview coming, it points at interview notes and real-world architectures with completed diagrams. If the goal is a bigger team, it points at the organization section.

That structure has a cost. There is no alphabetical index and no conceptual map, so it works when you already know what is broken and offers nothing to someone deciding what to learn first.

Entries are bare titles, so the curation is the whole contribution

Look at what an entry actually is. The Principle section runs from a paper on lessons from giant-scale services by Eric Brewer through Jeff Dean's designs and lessons from building large distributed systems, Joshua Bloch on API design, James Hamilton on efficiency and reliability and scaling, the Twelve-Factor App, the CAP theorem, ACID and BASE, consistent hashing, eventually consistent, cache is king, and understand latency. Each line is a title and a URL. There is no summary, no author note, no publication year, and no comment on when the advice stopped being current.

So the value of the repository is entirely in which links were chosen and where they were filed. That is a real contribution and it is not nothing, since nobody would assemble this order by accident. But it also means every claim in the list has to be checked at the source, and the described promise of an updated and organised list is a promise about ordering rather than about content.

A large share of the Principle links carry dates from 2007 to 2017

The URLs themselves are the evidence for how old this shelf is. Consistent hashing points at a 2007/11 path. Eventually consistent is a 2008/12 post. A High Scalability piece on database isolation levels is a 2011/2/10 path, and two more on that site are 2012/5/16 and 2014/5/12. Clean Architecture is a 2012/8/13 path. The probability of data loss in large clusters is 2017/01/26. A data access reading sits under a docs.microsoft.com path that begins with previous-versions, which is the host naming its own archive.

None of that makes a link wrong. Blocking versus non-blocking I/O, isolation levels and consistent hashing are not perishable, and a 2007 post on consistent hashing is still a good post. The consequence is narrower and worth stating: a link's presence here is not evidence that the practice is current, and for anything that has moved, the reader has to establish the vintage themselves.

Dead links are the reader's job, by the project's own account

There is a stated maintenance model and it depends on traffic. The community section welcomes contributions, points at the contribution guidelines in the repository root, and asks readers who find a link that is no longer maintained, or that is not a good fit, to submit a pull request. Read carefully, that sentence makes link rot a report from a reader rather than a check run by the project.

Nothing in the repository describes an automated link checker, a scheduled job, or a per-entry date, and the root holds only a handful of files. So the freshness of a section is a function of how many people happen to be reading that section, and a link in an unloved corner can sit dead for years. There is a pull-request template culture to its credit, and there is also no way to ask the repository which entries were verified this year. For a list whose entire claim is curation, that gap is the thing to hold in mind.

One section names its audience, the other ten do not

The intelligence section is the only one with a stated reader. The text says it is created for those who work with data and machine learning at big data and deep learning scale, and the other ten sections, from Principle through Talk and Book, name nobody. Most of them do not need to, since availability and stability are self-selecting. The organization section is the exception worth flagging: its stated subject is how technology companies increase team output and value rather than team size, and it is organised around hiring, management, organization, culture and communication.

That places career and management advice in the same table of contents as the CAP theorem, under the same curation, with no stated criteria for either. A reader who arrived for a database decision and clicks through to the organization section is being asked to take someone's hiring and communication writing as seriously as a consistency model, and the list offers no way to tell the difference.

The repository and the website it links are separate artifacts

The README opens by pointing at awesome-scalability.com, and that is the only homepage-shaped address in the whole project. The repository records none, which means the site is not the repository. A _config.yml in the root is the site generator configuration, so the pages are almost certainly generated from this content and then served from somewhere else, and the root contains little else: the README, a code of conduct, a contribution guide, a security policy, a licence file and a logo.

Two consequences. A fork gives you the content without giving you the site, and nothing in the tree says how the site being linked to is built or deployed. And the address in the README is plain http, so the canonical link a reader copies is not the secure one. For a project whose value is entirely in pointing at other people's writing, a pointer that can be swapped by whoever controls the redirect is worth replacing with the direct addresses before you rely on it.

No release, no version, and a last push on 2026-01-04

The maintenance facts are thin, and for a reading list that is not automatically a problem, since nothing here needs to be rebuilt. The repository is not archived, the last push was on 2026-01-04, and it publishes no GitHub releases. Its primary language is recorded as unknown, which is consistent with a repository that is one Markdown file, and the licence is MIT with the file in the root.

The gap that matters is not the last commit. It is that no entry anywhere carries a date, no tag exists to compare two states of the list, and a nine-month gap between pushes tells you nothing about whether the links inside were re-checked in that window. For most readers of a reading list that is acceptable, since the content is other people's content. For a team that wants to cite one of these entries in a design document, there is no provenance to cite, and the only defensible move is to open the source, check it yourself, and write down what you found and when.

Editorial conclusion

This list suits an engineer who already has a symptom and wants reputable places to read about it, and the symptom-based entry points are the part worth borrowing whatever else you decide. It does not suit someone building a curriculum, because entries are unannotated titles with no dates, and it does not suit anyone who has to know how fresh a link is, because nothing in the repository records when an entry was last checked. Read the Principle section first, and date what you find yourself relying on. MIT lets you fork the list, and a fork where you have added dates serves your team better than the original does.

Frequently asked questions

What is in the awesome-scalability reading list?

Eleven sections: Principle, Scalability, Availability, Stability, Performance, Intelligence, Architecture, Interview, Organization, Talk and Book. The README routes readers into them by situation, such as a system that is slow, one that is down, or an upcoming system design interview.

Does awesome-scalability check its links?

The project asks readers to submit a pull request when they find a link that is no longer maintained or is not a good fit, and points them at its contribution guidelines. No automated link check is described, and no entry carries a last-verified date.

How current are the links in awesome-scalability?

The last push was on 2026-01-04 and the repository publishes no releases. Many Principle entries point at URLs whose paths carry dates between 2007 and 2017, including one Microsoft documentation link under a previous-versions path.

What licence is awesome-scalability under?

MIT, with the licence file in the repository root. The root also carries a code of conduct, a contribution guide and a security policy.

Is awesome-scalability a repository or a website?

Both, and they are separate artifacts. The README links awesome-scalability.com in plain http, the repository records no homepage of its own, and a _config.yml site configuration sits in the root alongside a single Markdown file.

Official sources

  1. Official README
  2. Project repository