Inside twitter/the-algorithm: the For You feed as source code
GitHub describes it as Source code for the X Recommendation Algorithm. The repository metadata lists Scala as its primary language. The metadata lists the AGPL-3.0 license. This article stays within the project description and details documented in the GitHub repository README.
At a glance
- What is it?
- The AGPL-3.0 release covers the For You Timeline and Recommended Notifications, and roughly half of timeline posts still arrive from the search index. It ships no top-level Bazel workspace and no release tags.
- Who is it for?
- Use this repository as reading material, not as something to run. There is no top-level Bazel workspace, so there is no build or test cycle, and the last push was on 2025-09-08 with no release tags to pin.
- Can I use it commercially?
- Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
- Is it still maintained?
- Probably not. The repository last received commits 12 months ago, on September 8, 2025.
- What is it written in?
- Mainly Scala, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Half of the For You feed still comes from the search index
The For You Timeline is not one ranker. It is assembled from several candidate sources, and the one with the largest share is not a recommendation model at all. The search-index component finds and ranks In-Network posts, and roughly 50% of posts come from this candidate source. That ranking is done by light-ranker, the Light Ranker model used by the search index, which the repository calls Earlybird. In other words, the same retrieval and ranking stack that answers a search query supplies about half of the timeline a user scrolls. Everything else is out-of-network. tweet-mixer is the coordination layer for fetching Out-of-Network tweet candidates from underlying compute services, and UTEG, the user-tweet-entity-graph, keeps an in-memory User to Post interaction graph and finds candidates by traversing it. follow-recommendation-service supplies recommendations for accounts to follow and posts from those accounts. A reader who changes anything in the search ranking path is changing the timeline, whether or not they intended to.
The out-of-network path walks a graph kept in another repository
The out-of-network path has a structural problem for anyone auditing it: its engine is not here. UTEG is built on GraphJet, and GraphJet is a separate project at github.com/twitter/GraphJet. The code that decides which neighbours to walk is in this tree; the machinery that stores and traverses the graph is not. recos-injector sits on the same boundary, described as a streaming event processor for building input streams for GraphJet based services, so the events that drive the traversal come from a pipeline whose output you cannot read from here. graph-feature-service works one hop further out, serving graph features for a directed pair of users, with the example given being how many of User A's following liked posts from User B. The consequence is concrete: you can see that the graph is consulted and what shape the features take, but you cannot follow a single candidate post from event to ranked result without leaving the repository twice.
Two ranker stages, then home-mixer assembles the page
Ranking is a two-stage funnel, and the timeline and the notification path each run their own copy. On the For You path, light-ranker scores posts inside the search index, and heavy-ranker is the neural network that ranks candidates once sourcing is finished, described as one of the main signals used to select timeline posts. home-mixer is the main service used to construct and serve the Home Timeline, and it is built on product-mixer, the framework for building feeds of content. timelineranker sits alongside as a legacy service supplying relevance-scored posts from the Earlybird Search Index and the UTEG service, which tells you the older path is still in the tree. pushservice runs its own pair for Recommended Notifications. Its light ranker bridges candidate generation and heavy ranking by pre-selecting highly relevant candidates from the initial huge candidate pool, and its heavy ranker is a multi-task learning model predicting the probabilities that a target user will open and engage with a sent notification. Open and engage are separate predicted outputs, so that model optimises two targets at once.
visibility-filters mixes legal removal with revenue protection
One component carries several different reasons for suppressing a post, and that is where this repository stops being a clean description of moderation. visibility-filters is responsible for filtering X content to support legal compliance, improve product quality, increase user trust, and protect revenue. It does that through hard-filtering, visible product treatments, and coarse-grained downranking. Legal removal and revenue protection are therefore two inputs to the same subsystem, and the code gives you no way to tell them apart from the outside. Detection sits upstream in trust-and-safety-models, which holds models for detecting NSFW or abusive content. What you cannot get from this code is a per-post account of why something was suppressed: a hard filter for a legal request, a visible treatment, and a downrank all end with a post that is not in your timeline, and only the first is a removal in the ordinary sense. A reader auditing content suppression finishes with the mechanism and none of the policy.
Two surfaces ship, so Search and Explore stay closed
The scope is narrower than the opening line suggests. That line describes services and jobs serving feeds across all X product surfaces, naming the For You Timeline, Search, Explore and Notifications. The architecture section then states which product surfaces are included, and the list has two entries: the For You Timeline and Recommended Notifications. Search ranking, the Explore tab and trending logic are absent. Two more pieces a reader would expect are also absent, and both live in the separate the-algorithm-ml repository: TwHIN, the dense knowledge graph embeddings for Users and Posts, and heavy-ranker, whose role is described in the tables but whose implementation is not in this tree. The practical effect is that this repository cannot answer why a post ranked where it did in Search, and it cannot be used to study the embedding models, because following those table entries takes you out of the project.
No top-level workspace file, so nothing builds as one unit
There is no way to build this as a single thing. Bazel BUILD files are included for most components, but not a top-level BUILD or WORKSPACE file, and the project says a more complete build and test system is planned for the future. That one missing file is the difference between reading a component and verifying it. Without a top-level workspace you cannot resolve the dependency graph across services, and there is no test entry point to run, so nothing here can be compiled or exercised from a clean checkout. The tree is also broader than the architecture tables: ann/, ci/, cr-mixer/, docs/, science/ and simclusters-ann/ all appear as top-level entries without being described, next to a top-level RETREIVAL_SIGNALS.md that no table mentions. Read the tables as a guide rather than a map. cr-mixer has no row in any of the three component tables, so what it does is not described anywhere in the repository.
Last push was 2025-09-08 with no releases to pin against
For anyone pinning a study, the dates matter more than usual here. The repository is not archived, and the last push was on 2025-09-08. There are no GitHub releases, so there is no tag to point a reader at and no changelog explaining what changed. The default branch is main, which means a commit hash is the only stable reference, and a diff against main taken months from now will not match the code you read. Contributions are invited through GitHub issues and pull requests, and the project describes work on tools to manage those suggestions and sync changes to its internal repository. Security problems are steered away from public issues entirely, going to the official bug bounty program on HackerOne at hackerone.com/x. An engineer reproducing a ranking behaviour should record the commit hash and the Scala and Java source they read, because nothing in the repository will date the snapshot for them.
Editorial conclusion
Use this repository as reading material, not as something to run. There is no top-level Bazel workspace, so there is no build or test cycle, and the last push was on 2025-09-08 with no release tags to pin. Engineers who want to know how the For You feed is assembled will find the candidate sources, the two ranker stages and the filtering stage laid out plainly. Anyone expecting the Search or Explore ranking, or a runnable stack, should look elsewhere and pin a commit hash rather than trusting the main branch.
Frequently asked questions
What does twitter/the-algorithm actually contain?
It is the Scala and Java source for the services and jobs behind the For You Timeline and Recommended Notifications: tweetypie for reading and writing post data, home-mixer for assembling the feed, product-mixer as the feed framework, navi for model serving in Rust, and tweepcred for user reputation as a PageRank calculation.
Can you build and run the recommendation stack from this repository?
No. Bazel BUILD files are present for most components, but there is no top-level BUILD or WORKSPACE file, and a more complete build and test system is planned for the future. Read the code rather than trying to run it.
Which X surfaces are covered by the released code?
Only the For You Timeline and Recommended Notifications. Search and Explore are named as surfaces served by the same set of services, but their ranking code is not included in the repository.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/twitter-the-algorithm)