Open-source project
tairov/llama2.mojo avatar
tairov/llama2.mojo

llama2.mojo: Llama 2 Inference in a Single Mojo File

Inference Llama 2 in one file of pure 🔥

2,129 stars139 forksMojoMIT

At a glance

What is it?
llama2.mojo ports Karpathy's llama2.c to Mojo, using SIMD and parallelisation primitives to reach speeds above the original C implementation on Apple M1 Max hardware. It runs the stories and TinyLlama model families from a single .mojo source file and targets Mojo 1.1.0 with MAX 26.6.0.
Who is it for?
Developers learning Mojo who want a working example of SIMD vectorisation and parallel inference at the language level will find llama2.mojo a directly readable reference. The benchmark data in the README shows that on Apple M1 Max, the Mojo implementation outperforms the original llama2.c on the story models at the tested thread counts.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Mojo, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What llama2.mojo Is and Why It Was Written

Andrej Karpathy's llama2.c is a minimal C implementation of Llama 2 inference for the stories models: small educational models trained on children's stories data, ranging from 260K to 110M parameters. It serves as a readable reference implementation of transformer inference.

llama2.mojo is a port of that same implementation to Mojo, a systems programming language designed by Modular that exposes SIMD and vectorisation primitives directly to the programmer. The author started from a Python port (llama2.py) and moved it to Mojo, replacing Python's runtime with Mojo's compiled execution and hardware-level acceleration.

The stated motivation is to demonstrate Mojo's potential for hardware-level optimisations in a complete, running application. The result is a single source file, llama2.mojo, that implements the full inference pipeline. The tool targets developers who are learning Mojo, exploring its performance primitives, or comparing hardware-level languages for machine learning inference workloads.

Benchmark Results from the README

The README contains benchmark tables from a run dated 2026-09-12 using Mojo 1.0.0 on an Apple M1 Max with 8 workers. On the stories15M.bin model, llama2.mojo achieved 1283 tokens per second compared to 730 for llama2.c (parallelised) and 890 for llama.cpp (CPU, 6 threads). On stories42M.bin, llama2.mojo reached 506 tokens per second versus 270 for llama2.c and 420 for llama.cpp. On stories110M.bin, the Mojo implementation achieved 218 tokens per second against 102 for llama2.c and 187 for llama.cpp.

On TinyLlama-1.1B, the README shows 23 tokens per second for llama2.mojo. No comparison figure for llama2.c is given for that model on M1 Max.

On an Ubuntu x86 AVX-512 VM with 4 vCPU (Intel Xeon Skylake), llama2.mojo with 4 workers reached 327 tokens per second on stories15M.bin versus 292 for llama2.c with OpenMP. On the same platform, single-threaded llama2.mojo achieved 154 tokens per second against 105 for single-threaded llama2.c.

These figures come from the README and reflect specific hardware and Mojo version combinations. The README labels them as benchmark measurements, not as general performance guarantees.

Installing Mojo and Running the First Inference

The README documents the Mojo 1.1.0 installation using the uv package manager. The old modular package was retired in Modular 26.6. Install the new way:

bash
uv venv --python 3.13 ~/.modular-venv
uv pip install --python ~/.modular-venv/bin/python mojo "max[all]"
source ~/.modular-venv/bin/activate
mojo --version   # Mojo 1.1.0

Linking a Mojo executable also requires a C compiler on the host (apt-get install gcc on Debian or Ubuntu). Once Mojo is installed, clone the repository and navigate into it:

bash
git clone https://github.com/tairov/llama2.mojo.git
cd llama2.mojo

Download a model file:

bash
wget https://huggingface.co/karpathy/tinyllamas/resolve/main/stories15M.bin

Run inference with a prompt:

bash
mojo llama2.mojo stories15M.bin -s 100 -n 256 -t 0.5 -i "Once upon a time"

The README's example output shows the model generating a children's story continuation from the prompt.

Command-Line Options and Supported Models

llama2.mojo accepts several command-line flags documented in the README:

- -s sets the random seed (default: current time in milliseconds) - -n sets the number of steps to run (default: 256; 0 uses max_seq_len) - -t sets temperature in the 0 to 1.0 range (default: 0.9) - -i provides an input prompt string - -z sets the tokenizer path (default: tokenizer.bin) - -j sets the number of parallel workers, defaulting to the number of performance cores - -pc controls whether the config is printed (0 or 1)

The supported models as listed in the README are the stories family (260K, 15M, 42M, 110M parameters) and TinyLlama-1.1B-Chat-v0.2. These are the models that have been verified to run via llama2.mojo. Larger Llama 2 models are not listed as supported.

A tokenizer.bin file is included in the repository for use with the stories models. The TinyLlama model uses the same tokenizer format and can be downloaded from HuggingFace.

How Mojo's SIMD Primitives Enable the Performance

The performance advantage of llama2.mojo over llama2.c on the tested hardware comes from Mojo's SIMD and vectorisation features. Mojo allows the programmer to express vectorised operations directly, similar to writing SIMD intrinsics in C but with a higher-level syntax that the compiler can map to hardware vector units.

The README credits the vectorisation of matrix multiplication as the key optimisation. Standard llama2.c uses auto-vectorisation hints to the C compiler; Mojo's SIMD primitives give more explicit control. The README's mention of a persistent worker pool build (PR #102) in the 2026-09-12 benchmark run indicates that parallelisation across multiple hardware threads is also a factor in the multi-core results.

The go.mod is not relevant here; this is a pure Mojo project. The Dockerfile in the repository uses the older modular install method and is documented as targeting an older SDK version, reflecting the project's history before the modular package retirement.

Limitations: Model Scope, Platform, and Production Use

llama2.mojo is explicitly limited to the stories and TinyLlama model families. It does not support full Llama 2 7B, 13B, or 70B models, nor any models outside the llama2.c checkpoint format. Users who need to run production LLMs at scale should use llama.cpp, Ollama, or a framework designed for that purpose.

GPU inference is not documented. llama2.mojo runs on CPU. The parallelisation is over hardware cores via Mojo's parallelize primitive. For GPU inference, dedicated frameworks are necessary.

Mojo is a relatively new language that has undergone significant API changes. The README specifies Mojo 1.1.0 as the supported version (with MAX 26.6.0). Earlier versions used the modular package for installation, which has since been retired. Developers who have an older Mojo setup must reinstall using the uv path. The upgradelog.md file in the repository documents the changes required across Mojo version upgrades.

Repository Layout and Citing the Project

The top-level layout is minimal, reflecting the single-file goal. The main source is llama2.mojo. The repository also contains tokenizer.bin for the stories model, a t260.bin test model, a tests/ directory, a Dockerfile for containerised runs, a gradio_app.py providing a Gradio-based web interface, and agents-optimization.md documenting agent-level performance notes.

The README includes a BibTeX citation block for academic use. Researchers who discuss or build on llama2.mojo in published work are asked to cite the project.

The project is licensed under MIT, which permits commercial use, modification, and distribution. It is not archived; the last push was on 2026-09-20. There are no GitHub releases; the project is distributed directly from the repository.

Editorial conclusion

Developers learning Mojo who want a working example of SIMD vectorisation and parallel inference at the language level will find llama2.mojo a directly readable reference. The benchmark data in the README shows that on Apple M1 Max, the Mojo implementation outperforms the original llama2.c on the story models at the tested thread counts. Users who need GPU inference, support for larger models, or a production-ready inference server should use llama.cpp, Ollama, or another framework. Before running, confirm Mojo 1.1.0 is installed via the uv path documented in the README, since the old modular package was retired in Modular 26.6.

Frequently asked questions

What models does llama2.mojo support?

llama2.mojo supports the stories model family (260K, 15M, 42M, and 110M parameters) and TinyLlama-1.1B-Chat-v0.2. These are checkpoint-format models compatible with the llama2.c format. Larger Llama 2 models and other architectures are not listed as supported.

How do you install Mojo to run llama2.mojo?

Use uv to create a Python 3.13 virtual environment and install the mojo and max packages. The README's install commands are: uv venv --python 3.13 ~/.modular-venv, then uv pip install mojo max[all] into that environment. The old modular package was retired in Modular 26.6.

Does llama2.mojo support GPU inference?

The README does not document GPU inference. llama2.mojo runs on CPU using Mojo's SIMD and parallelisation primitives. The -j flag controls the number of parallel worker threads, defaulting to the count of performance cores.

Official sources

  1. Issues
  2. License: MIT
  3. Project website
  4. README
  5. tairov/llama2.mojo on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/tairov-llama2-mojo.svg)](https://hysenlabs.com/projects/tairov-llama2-mojo)