Open-source project
dendibakh/perf-ninja avatar
dendibakh/perf-ninja

dendibakh/perf-ninja: a lab-based course for low-level performance tuning in C++

This is an online course where you can learn and master the skill of low-level performance analysis and tuning.

3,878 stars408 forksC++License varies

At a glance

What is it?
Performance Ninja teaches CPU cache misses, branch mispredictions and dependency chains through C++ lab assignments that are graded by automated benchmarking on GitHub. It is a practice-first course, not a library, and the README lists it as work in progress.
Who is it for?
Adopt Performance Ninja if you already write C++ and want to reason about cache misses, branch mispredictions and dependency chains on real hardware, and you accept that the course is described as work in progress with empty CPU Frontend Bound and Data-Driven optimization categories. Skip it if you need a library to drop into a product, or if you cannot read assembly, since the README lists that as a plus and the labs are graded on measured speedups.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What problem Performance Ninja solves, and for whom

Most performance advice stops at the level of a profiler screenshot. Performance Ninja goes one level down: the README describes it as an online course where you learn to find and fix low-level performance issues, for example CPU cache misses and branch mispredictions. The delivery format is lab assignments plus YouTube videos, and the README is explicit that you spend at least 90 percent of the time analyzing performance of the code and trying to improve it.

The intended audience is narrow. Basic C++ skills are described as an absolute must-have, and knowledge of compilers, computer architecture and the ability to read assembly code is listed as a plus. Denis Bakhvalov's book "Performance Analysis and Tuning on Modern CPUs" is recommended as an introduction to the basics, which tells you the course assumes you have already met the vocabulary of IPC, cache levels and branch predictors.

That narrowness is the point. A general backend engineer optimizing a database query will find little here. Someone who wants to know why a tight loop runs at 1.2 instructions per cycle instead of 3 will find a structured path, because each lab isolates one effect and asks you to change the code until the measured number moves.

How the lab assignments and automated benchmarking fit together

The repository is organized by bottleneck category rather than by difficulty. Labs live under labs/ in directories named core_bound, memory_bound, bad_speculation, misc, and two categories, CPU Frontend Bound and Data-Driven optimizations, are listed in the README with no entries yet. The named labs are concrete: Vectorization 1 and 2, Function Inlining, Dependency Chains 1 and 2, four Compiler Intrinsics labs, Data Packing, Loop Interchange 1 and 2, Loop Tiling, SW memory prefetching, False Sharing, Huge Pages, Memory Order Violation, Memory Alignment, Branches To CMOVs, Conditional Store, Replacing Branches With Lookup Tables, C++ Virtual Calls, plus Warmup, LTO, PGO and Optimize IO under Misc.

The workflow the README describes is: read the assignment, improve the code, then submit your solution to GitHub for automated benchmarking and verification. The buildbot/ directory and the CI workflow badges at the top of the README are the machinery behind that sentence. This is the design decision that separates Performance Ninja from a tutorial repository. A tutorial can tell you that your version is faster; this one measures it on fixed hardware and rejects a solution that does not actually improve the benchmark.

The cost of that design is that grading is tied to the CI runners. The README says the course is supported on Linux, Windows and Mac and runs on recent hardware including Intel 12th-gen Alderlake, AMD Zen3 and Apple M1, and the workflow badges name CI_Linux_Alderlake, CI_Macos_M1, CI_Win_Zen3 and CI_Linux_Coffeelake. Four named configurations is a real constraint: if your machine is a different microarchitecture, the absolute numbers you see locally will not match what the grader sees, and a change that helps on Zen3 may do nothing on Alderlake.

Installing Performance Ninja and running the warmup lab

There is no package to install. The course is the repository, and the README points you at a Get Started page before you touch a lab: "Before you start working on lab assignments, make sure you read Get Started page and watch the warmup video." The top level of the repository also carries QuickstartLinux.md, QuickstartMacOS.md and QuickstartWindows.md, so the platform-specific setup steps live there rather than in the README. Clone the repository first:

bash
git clone https://github.com/dendibakh/perf-ninja.git
cd perf-ninja

Then read the quickstart for your platform and the general Get Started page, which is where the toolchain requirements and the submission flow are documented. The README does not restate the compiler versions or the benchmark harness flags, so do not guess them from the lab directories.

The first lab to run is the warmup, which the README lists under Misc as labs/misc/warmup. Its purpose is to prove the pipeline works end to end before you spend four hours on a memory-bound assignment:

bash
ls labs/misc/warmup

What you should see is the lab's own source and configuration files. The README does not document the exact build command for a single lab, so follow the quickstart file for your OS rather than a command copied from a blog post. Once the warmup builds and produces a benchmark result, you have the same loop the rest of the course uses: edit, measure, submit.

Where Performance Ninja is the wrong tool

The README states plainly that the project "is in a very much work-in-progress state" and that new lab assignments and videos will be added. Two of the five bottleneck categories, CPU Frontend Bound and Data-Driven optimizations, appear in the README with no labs listed under them. If your goal is a complete curriculum that covers the whole Top-Down analysis tree in order, the repository does not currently offer that, and no release history is available to tell you how often labs arrive.

The second limitation is the grading model. Because solutions are verified by automated benchmarking, you are optimizing against a measurement on specific CI hardware, not against a universal truth. A change that shortens a dependency chain on one microarchitecture can be neutral on another, and the README's own list of supported machines is the boundary of what has been validated. If you work on an unusual target such as an embedded core or a GPU, the labs will teach you the reasoning but the numbers will not transfer.

Finally, the prerequisites are real filters. The README calls basic C++ an absolute must-have and recommends a book as preparation. A Python developer who wants to understand why numpy is fast will spend most of the course fighting the language rather than the bottleneck. The Rust and Zig ports linked from the README exist, but they are separate repositories maintained by other people, so their state is not covered by anything in this one.

How it compares with reading the book or a static tutorial

The obvious alternative is the recommended book itself, "Performance Analysis and Tuning on Modern CPUs". The difference in approach is measurement versus explanation. A book can describe a false-sharing cache line and draw the two-thread diagram, but it cannot tell you whether your fix reduced the benchmark time on the machine in front of you. Performance Ninja inverts that: the README says you will spend at least 90 percent of the time analyzing performance of the code, and the automated verification means a wrong mental model produces a failing submission rather than a vague feeling of improvement.

Against a typical tutorial repository, the difference is the grader. A tutorial shows a before-and-after snippet and asks you to trust the author's laptop. Here the CI workflows in .github run the benchmark, which is why the README can list specific CPU families as supported targets. That also means the course is heavier to set up than copying a snippet, and it is the reason the quickstart files exist per operating system.

A third option is a general profiler-driven course built around a vendor tool. Performance Ninja is not tool-specific in the way the README is written; it names the bottleneck categories and the labs, not a single profiler. For someone who wants to learn the reasoning that survives a change of tooling, that is the more durable framing.

Maintenance, licensing and the cost of upgrading

The repository is not archived and the last push was on 2026-09-15, so the codebase is being touched. There are no retrieved releases, which fits a course repository: there is no version to pin and no changelog to read before upgrading. In practice, upgrading means pulling main and rebuilding the labs, and the risk is that a lab's source or its benchmark configuration changed under you. If you are working through the course over weeks, that is worth knowing before you start a four-hour assignment.

Maintenance is funded rather than corporate. The README asks for support through GitHub Sponsors, Patreon or PayPal and lists current sponsors, and it says sponsorship speeds up adding new lab assignments. That is a direct statement about the project's development rate: the pace depends on contributions and sponsorship, not on a roadmap with dates.

The licence is the part to read carefully. The README ends with "Copyright © 2025 by Denis Bakhvalov under Creative Commons license (CC BY 4.0)", which is a content licence, not a software licence, and no separate licence file appears in the top-level repository entries. CC BY 4.0 is designed for creative works and is unusual for source code; if you intend to reuse lab sources in your own teaching material or product, check the terms with your own counsel rather than assuming a permissive software licence applies. This article gives no legal advice.

Editorial conclusion

Adopt Performance Ninja if you already write C++ and want to reason about cache misses, branch mispredictions and dependency chains on real hardware, and you accept that the course is described as work in progress with empty CPU Frontend Bound and Data-Driven optimization categories. Skip it if you need a library to drop into a product, or if you cannot read assembly, since the README lists that as a plus and the labs are graded on measured speedups. Before committing time, open GetStarted.md and the Quickstart file for your platform, run the warmup lab at labs/misc/warmup, and confirm the CI workflow for your CPU family exists in .github, because a lab you cannot get graded is a lab you cannot finish.

Frequently asked questions

Is Performance Ninja free to use?

The README says the course is free by default and asks users to support the project through GitHub Sponsors, Patreon or PayPal. It also states the material is copyright Denis Bakhvalov under a Creative Commons CC BY 4.0 licence.

What C++ skills do I need before starting Performance Ninja?

The README calls basic C++ skills an absolute must-have and recommends the book Performance Analysis and Tuning on Modern CPUs as an introduction. Knowledge of compilers, computer architecture and the ability to read assembly code is listed as a plus.

Which operating systems and CPUs does Performance Ninja support?

The README says the course is supported on Linux, Windows and Mac, and runs on recent hardware including Intel 12th-gen Alderlake, AMD Zen3 and Apple M1. The CI badges name Linux Alderlake, macOS M1, Windows Zen3 and Linux Coffeelake configurations.

Does Performance Ninja have Rust or Zig versions?

Yes. The README states the lab assignments are implemented in C++ and that the course was ported to Rust as perf-ninja-rs and to Zig as perf-ninja-zig, both maintained in separate repositories by other authors.

How are Performance Ninja lab solutions checked?

The README says that once you finish improving the code you submit your solution to GitHub for automated benchmarking and verification. The buildbot directory and the CI workflow badges are the parts of the repository that carry out that grading.

Official sources

  1. dendibakh/perf-ninja on GitHub
  2. Issues
  3. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/dendibakh-perf-ninja.svg)](https://hysenlabs.com/projects/dendibakh-perf-ninja)