# llama2.mojo benchmarks itself against llama2.c on three machines, and wins one table by one percent

> A Llama 2 inference implementation in a single Mojo file, ported from the author's Python version and from the well known C reference. The README's own tables are the most interesting part, because two of the three machines show a win that is real but small, and the third shows a loss.

**tairov/llama2.mojo** — Inference Llama 2 in one file of pure 🔥

- Repository: https://github.com/tairov/llama2.mojo
- Website: https://www.modular.com/blog/community-spotlight-how-i-built-llama2-by-aydyn-tairov
- Stars: 2,129 · Forks: 139
- Language: Mojo
- License: MIT
- Published: 2026-09-15 · Updated: 2026-09-15 · Language: en
- Canonical page: https://hysenlabs.com/projects/tairov-llama2-mojo

## The headline number is a thread-count comparison

The opening claim is that this port beats the C reference by thirty percent in multi-threaded inference and beats a widely used CPU inference tool by twenty percent on the smallest checkpoint.

Both figures are thread counts in disguise, and the README is honest enough to publish the single-thread columns next to the parallel ones.

On the virtual machine with four processors, the parallel column is roughly one and a quarter times the C reference on the small checkpoint, while the single-worker column is about one and a half times the single-threaded C reference. On a six-core desktop part, the parallel column is essentially a tie with the C reference on the small checkpoint, and the Mojo entry loses outright on the next model up.

So the thirty percent figure is the best case from one of three tables, not a general property. That is not a criticism of the author, who put the unflattering columns in the same table as the flattering ones. It is a reason to read the third table before repeating the first claim.

The single biggest number on the page is a different comparison: nearly two hundred and fifty times the Python version, which is what you would expect from moving to a compiled language at all.

## Every benchmark row is stamped with a date and a build

There is a detail here that costs the author nothing and tells you more than the tables do.

The benchmark section carries an italic note recording when it was last run, which language version was used, and which specific build configuration was measured, including a reference to the pull request that introduced it and a mention of a persistent worker pool.

That is more provenance than most benchmark tables in open source carry, and it matters here for a specific reason. A systems-language inference implementation is exactly the kind of project whose numbers move when the compiler changes, because the parallelism primitives are the product. A table without a date and a build is close to meaningless in that situation.

The page also links an external multi-language benchmark of the same project family on an Apple laptop, which is a cross-check by someone else rather than a second self-reported table.

Two of the three machine rows are virtual machines rather than dedicated hardware, which the page does state. Virtual-machine numbers are noisy in a way dedicated-metal numbers are not, so the margins on those rows carry more uncertainty than the single-digit margins would otherwise suggest.

## The supported model list is five checkpoints and nothing larger

The model table is short and it is the clearest scope statement in the repository.

Four checkpoint sizes from a small story-generation series, plus one chat-tuned model at about a billion parameters. That is the whole list.

Those checkpoints are deliberately tiny, which is why they exist: they produce grammatically plausible but semantically empty stories, as the example output shows, where a child picks up a knife and the sentence trails off mid-warning. That output is included verbatim in the README, truncated mid-sentence.

Including it is the right call. A reader evaluating this project needs to see that the output is not going to be useful for anything, and the author shows it rather than curating a better sample.

The billion-parameter chat model runs at twenty-three tokens per second with eight workers on the laptop in the table, which is the honest number for a model of that size on CPU. If you are evaluating this for real inference work, that row is the one to look at rather than the small-checkpoint rows, because the small ones are where the vectorisation wins are largest.

## The documented toolchain is current, the container is not

There are two install paths in this repository and only one of them looks maintained.

The documented path is precise: it names the supported language version and the maximum version that works, installs a virtual environment with a specific Python, installs the compiler and the parallelisation package from that environment, activates it, and prints the version to confirm. It also notes that an older package name was retired in a specific release, which is exactly the kind of note that saves an afternoon.

It adds one non-obvious requirement: linking the executable needs a C compiler on the host, with the distribution-specific package named.

The container definition is a different story. Its base image is a distribution release from 2020, it downloads a Miniconda installer by a pinned URL rather than by package, it installs a set of packages including an editor and a version control client that the project does not need, and it authenticates to the vendor with a build argument that defaults to a placeholder before piping an install script from a URL into a shell. It then calls the retired package name that the README says no longer exists.

So the container will not build from the current toolchain. The README's manual path is the one to use.

## The default worker count is clamped to physical cores

The command line options are listed with their defaults, and two of them are worth reading closely.

The worker count defaults to the number of performance cores rather than the number of logical ones, which is the right default for a vectorised workload where the execution units are shared. And the values are clamped to that count, so asking for more workers than you have performance cores does nothing.

```bash
mojo llama2.mojo stories15M.bin -s 100 -n 256 -t 0.5 -i "Once upon a time"
```

That clamping is quietly important when reading the benchmarks. It means the single-worker column is a genuine configuration rather than a hypothetical, and it means nobody can accidentally produce a number by oversubscribing.

The other options are unremarkable but complete: a random seed that defaults to the current time in milliseconds, so runs are not reproducible unless you pass one; a step count that defaults to a fixed number and where zero means use the model's maximum sequence length; a temperature with a default slightly above the middle of the range; a prompt; and a tokenizer path with a default filename.

A seed defaulting to wall-clock time is a design choice worth noticing. For a demo it is convenient, and for anything you want to reproduce you must remember to pass one.

## A model checkpoint and a tokenizer are committed to the repository

The root listing has two surprises in it for a code repository.

There is a binary model file at the top level, several hundred megabytes by the look of the example output's reported checkpoint size, and a tokenizer file beside it. So the repository carries weights, not just the code that loads them.

For a project whose quick start downloads a checkpoint from a model hub, that is redundant. The documented path fetches the small checkpoint over the network before running. Having a copy in the tree is convenient for anyone cloning without network access, and expensive for everyone who clones with it.

Alongside those there are the things a research single-file project actually needs. There is the Mojo source file itself, which is the product. There is a Python interface file, a shell script to run tests, a test directory, and an upgrade log. There is a document about optimising agent behaviour, which is unrelated to the inference code and appears to be a note carried over from another project.

The upgrade log is worth a look when evaluating maintenance. A language that retires packages between releases, as this one did, will break a project like this, and a changelog is how you would find out.

## Citation is requested for a research port of a research reference

The closing section asks you to cite the project if you use or discuss it in academic research, and supplies a BibTeX entry.

The request is slightly unusual in this context, and the reason is worth naming. This is a port of an existing educational reference implementation, and the chain of attribution has three links: the original C implementation the author did not write, the author's own Python port which is the direct parent of this code, and the language runtime whose vectorisation primitives produce the speedup.

Only the middle link gets a citation request. The first is credited throughout the README by link in every benchmark table header. The third is credited in the prerequisites and the container file's copyright header, which carries the vendor's own licence and copyright notice intact, which is more than a link.

So the attribution is present at every level; the formal request is just narrower than it might appear.

For a reader weighing whether to build on this, the practical note is that the code is MIT licensed and small enough to read in an afternoon, which makes it a better teaching artefact than a dependency.

## Conclusion

This is a good project to read if you want to see what a systems language buys you on CPU inference, and a poor one to lift into production. It targets the smallest Llama checkpoints in a research series and generates visibly weak prose, which is the point of those checkpoints but not what you would serve. Two cautions before you copy the approach. The headline comparison depends on how many worker threads each implementation is given, and the README is careful enough to publish a single-worker column that shows the margin shrinking dramatically. And the container definition in the repository is built around a retired toolchain, so treat the documented install path as the supported one.

## FAQ

### What is llama2.mojo?

It is a Llama 2 inference implementation written in a single Mojo file, ported by the author from their own Python version after the language's release. It uses the language's vectorisation primitives to run the transformer forward pass on CPU, with a persistent pool of parallel workers and a batched matrix multiplication path.

### How much faster is llama2.mojo than the C reference?

It depends on the machine and the worker count, and the README publishes both. On a four processor virtual machine the parallel column is roughly a quarter faster on the smallest checkpoint. On a six-core desktop part the parallel column is a tie on that checkpoint and behind on the next model up. The thirty percent headline is the best case from one of three tables, with single-worker columns shown alongside.

### Which models does llama2.mojo support?

Four checkpoint sizes from a small story-generation series, up to about 110 million parameters, plus one chat-tuned model at roughly 1.1 billion. That is the complete list, and the page includes example output showing the small checkpoints produce grammatically plausible but empty prose.

### How do I install the Mojo toolchain for llama2.mojo?

Create a virtual environment with a specific Python version, install the compiler and the parallelisation package into it from that environment, then activate it and check the reported version. The page names the supported language version and the maximum that works, notes that an older package name was retired, and says a C compiler is needed on the host to link the executable.

### Does llama2.mojo support multiple worker threads?

Yes. A worker count option defaults to the number of performance cores and is clamped to that count, so oversubscribing does nothing. Benchmarks are published with both parallel and single-worker columns, and the single-worker figures show how much of the advantage comes from threading.

## Sources

- [Issues](https://github.com/tairov/llama2.mojo/issues)
- [License: MIT](https://github.com/tairov/llama2.mojo/blob/master/LICENSE)
- [Project website](https://www.modular.com/blog/community-spotlight-how-i-built-llama2-by-aydyn-tairov)
- [README](https://github.com/tairov/llama2.mojo/blob/master/README.md)
- [tairov/llama2.mojo on GitHub](https://github.com/tairov/llama2.mojo)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/tairov-llama2-mojo
