are-we-learning-yet is a catalog with a Rust scraper behind it, and the ordering score is undocumented
How ready is Rust for Machine Learning?
At a glance
- What is it?
- A static site that answers whether Rust is ready for machine learning by listing the crates, published to GitHub Pages on every merge to master. The interesting parts are the seams: a scraper that must be rerun by hand because the dev server will not trigger it, a weekly cron that leaves download counts up to a week stale, and a Creative Commons licence covering Rust source as well as prose.
- Who is it for?
- Use it as a starting index rather than a dependency list, and read the ordering score as an editorial convenience rather than a quality measure, since the repository does not say what goes into it. If you want to contribute, the real work is curation in the content pages plus a scraper run, and remember that the dev server will not pick up a change to crates.yaml on its own.
- Can I use it commercially?
- Yes, with credit. CC-BY-4.0 allows commercial use as long as you credit the authors and indicate what you changed. It is written for creative content, so check how it applies to any code.
- Is it still maintained?
- Yes. The repository last received commits 5 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
It is a static site with a scraper bolted to one side, not a library
The repository is two things that share a GitHub Pages deployment. One is a scraper, written in Rust, that reads a hand-maintained list at _data/crates.yaml, pulls additional metadata from crates.io and the GitHub API, generates a score used for ordering crates, and writes the result to _data/crates_generated.yaml. The other is the site content, which follows the layout established by cobalt.rs and is built into a _site directory by cobalt build. Cobalt is the static site generator, the minimum version is 0.20.0, and CI builds with 0.20.4. There is no Rust toolchain file at the top level, so the stated floor of Rust 1.85 or newer is prose in the README rather than something the repository enforces. Nothing here is importable; the value is the catalog and the maintenance apparatus around it.
The score that orders crates is never explained
The scraper exists to produce one number per crate, described only as a score used for ordering crates. The README does not say what goes into it: not whether download counts dominate, not whether recency is weighted, not whether a crate with a higher score is considered more mature or more relevant to machine learning. That matters because the site inherits the answer. Anyone reading the catalog sees an order and will read it as a judgement, when the mechanism behind the judgement is in the scraper source under _scraper and in the YAML that seeds it. The companion content pages are the manual half of the same catalog, covering what the project is, how entries are chosen, and the timeline, which is a reasonable division of labour only if you know where the boundary sits. It is worth reading the scraper before trusting the order it produces.
cobalt serve will not rerun the scraper when crates.yaml changes
The dev server rebuilds the site whenever content changes, and it is easy to read that as everything regenerating. It is not. cobalt serve does not trigger the scraper when crates.yaml changes, nor when the scraper itself changes, so the file that drives the whole catalog can be edited, served locally, and appear to work while the generated data still reflects the previous run. The stated remedy is to rerun just scrape to update crates_generated.yaml, which makes the edit-test loop two commands rather than one. That is the sharpest seam in the project: the stale state is silent, because the site renders. The cache makes it worse in a second way. Fetched data is stored in a _tmp directory to speed up repeated generation and avoid hammering the APIs, so a fix in the upstream metadata does not show up either. just clean drops both the generated data and the cached responses.
just serve pins port 3000 because cobalt serve alone does not
Two port behaviours sit side by side. Plain cobalt serve picks an arbitrary free port, which is convenient for one developer and useless for a link in a README or a second terminal window. just serve passes --port 3000, so the site lands on localhost:3000 predictably. The task runner is just, taken as a recommendation rather than a requirement, with the commands also written out in the Justfile. The other recipes are just build, which is scrape plus cobalt build, and just check, which is the same formatting and lint gate CI runs. That last one is the piece to wire into a pre-commit hook: the project has already decided what its own quality bar is, and it is expressed as a single recipe rather than a wall of flags. Clean is the third, and it is the one to reach for first when generated data looks wrong.
Without a GITHUB_TOKEN the GitHub API side returns 403
The scraper talks to two upstreams, and one of them needs credentials. Exporting GITHUB_TOKEN is what avoids 403 rate limiting errors while generating crate data, and the recommended way to supply it is a .env file, which the task runner picks up automatically. That is the whole setup story for local development:
export GITHUB_TOKEN=<YOUR_GITHUB_TOKEN>
# Scrape crate/repo data for sitegen
just scrape
# Start a dev server on port 3000
just serveThe stated requirement list is short on purpose, a Rust toolchain for the scraper, cobalt for the site, and just if you want the recipes. What the documentation does not discuss is what happens on a partial run, where crates.io metadata fetched fine and the GitHub half failed partway through, and whether the output is written atomically or left in an intermediate state. Given that the cache exists specifically to avoid abusing these APIs, that failure mode is the one worth checking yourself the first time the numbers look thin.
Publishing is automatic, but the statistics only move once a week
Deployment has two triggers and they are not the same. Every merge into master is published automatically by a GitHub Actions job, so a content fix goes live without anyone touching the site configuration, and the custom domain is a CNAME file at the root with the workflow living under .github. Separately, a weekly cron job runs the publishing task as well, and its stated purpose is keeping crate statistics current: download counts and stars. That job can also be triggered by hand from the Actions tab with workflow_dispatch, which is the escape hatch when you know a number is wrong and cannot wait a week. The repository publishes no tagged releases at all, so there is no version to point at. For a catalog whose selling point is current ecosystem state, that combination, no tags on the one hand and weekly numbers on the other, is the detail to keep in mind when you cite a figure.
CC-BY-4.0 covers the Rust source as well as the writing
The licence is Creative Commons Attribution 4.0, which is a content licence rather than a software licence, and it covers the whole repository, including the scraper that fetches data and the Justfile that drives it. For a documentation site that is an unremarkable choice, since the prose is the product and attribution is a reasonable ask. For the Rust code it is an unusual one, because CC-BY-4.0 is not on the list of licences that most package ecosystems and corporate scanners accept for dependencies, and it carries no warranty or patent grant of the kind an Apache or MIT licence provides. There is also a small inconsistency inside the README itself: the same site is linked once as an absolute address over plain http in the opening paragraph and once as a bare relative path in the publishing section, so the two links do not resolve the same way. Neither point blocks using the catalog, but both are worth raising with the maintainer if you want to depend on the code.
The hand-written pages and the generated ordering have to agree
Because curation is manual and ordering is generated, the site can disagree with itself. The content pages describe what belongs in a Rust machine learning catalog and when things arrived; the scraper decides the order of the entries on the page. Neither mentions the other, so a crate can be described accurately in prose and sit in a position that implies the opposite judgement. That is not a bug so much as an architectural consequence of splitting the responsibility this way, and it is the same reason the project asks for pull requests in two categories: adding missing crates, and improving the content. Contributors who only do the first get a new row; contributors who do the second decide what the row means. Feedback, issues, and pull requests are all invited for both. When you are reading the site, the prose is the part with a human behind it, and the ordering is the part with a scraper behind it.
Editorial conclusion
Use it as a starting index rather than a dependency list, and read the ordering score as an editorial convenience rather than a quality measure, since the repository does not say what goes into it. If you want to contribute, the real work is curation in the content pages plus a scraper run, and remember that the dev server will not pick up a change to crates.yaml on its own. Before you cite a number from the site, check how old the download counts are, because the refresh is a weekly cron rather than a per-merge step.
Frequently asked questions
What is are-we-learning-yet?
A catalog of the Rust machine learning ecosystem, published as a static site at arewelearningyet.com and inspired by Are We Web Yet. The repository is not a library: it is site content plus a scraper that generates the crate data.
How does are-we-learning-yet build its crate list?
The scraper reads _data/crates.yaml, fetches metadata from crates.io and the GitHub API, generates a score used for ordering crates, and writes _data/crates_generated.yaml. Fetched data is cached in a _tmp directory, and just clean drops both the generated data and the cached responses.
Do I need Rust to contribute to are-we-learning-yet?
Only for the scraper, which needs a Rust toolchain 1.85 or newer. Site generation needs cobalt 0.20.0 or newer, with CI building against 0.20.4, and just is recommended as the task runner for the scrape, serve, build, check, and clean recipes.
Why do I need a GITHUB_TOKEN for are-we-learning-yet?
Exporting GITHUB_TOKEN avoids 403 rate limiting errors while the scraper generates crate data, and the token can be set in a .env file that the task runner picks up automatically. After that, just scrape regenerates the data and just serve starts the dev server on port 3000.
How current are the numbers on are-we-learning-yet?
Every merge into master publishes automatically through a GitHub Actions job, while download counts and stars are refreshed by a separate weekly cron run that can also be triggered by hand with workflow_dispatch. The repository publishes no tagged releases, so there is no version to compare against.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/anowell-are-we-learning-yet)