Open-source project
xai-org/grok-1 avatar
xai-org/grok-1

grok-1: running the 314B open weights release with JAX

GitHub describes it as Grok open release. The repository metadata lists Python as its primary language. The metadata lists the Apache-2.0 license. This article stays within the project description and details documented in the GitHub repository README.

52,240 stars8,545 forksPythonApache-2.0

At a glance

What is it?
The xai-org/grok-1 repository is JAX example code for loading and sampling from the Grok-1 open weights, not a serving stack. Two files do the work, the model layer was written for correctness rather than speed, and the dependency pins choose a CUDA 12 build for you.
Who is it for?
Use this repository when your goal is to confirm the Grok-1 weights load and generate, to read a compact reference for the architecture, or to start a serving stack you will write yourself. Do not adopt it as a running product: there is no server, no CLI surface, no test suite in the tree, and the MoE layer is explicitly not written for throughput.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Probably not. The repository last received commits 25 months ago, on August 30, 2024.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

314B parameters with two of eight experts active per token

The architecture is specified precisely enough to plan against, and the specification is the most useful thing in the repository. Grok-1 has 314B parameters arranged as a Mixture of 8 Experts with 2 experts used per token, across 64 layers, with 48 attention heads for queries and 8 for keys and values, and an embedding size of 6,144. Tokenization is a SentencePiece tokenizer with 131,072 tokens, position information comes from rotary embeddings, and the maximum sequence length is 8,192 tokens. The last number is the one that constrains a product decision before hardware does. Activating two experts per token is the lever that makes a mixture of experts cheaper than a dense model of the same size, and the README still states that a machine with enough GPU memory is required to test the code with the example. Nothing here is a small model you can try on a laptop.

run.py samples one test input and exits, so reading it is the tutorial

The whole first run is two commands, after the weights are in place.

shell
pip install -r requirements.txt
python run.py

The script loads the checkpoint and samples from the model on a test input, which means the repository's entire user interface is a script with a hardcoded prompt and a hardcoded sampling call. There is no server to start, no flag to pass a prompt, and no streaming output in the code that ships here. To ask something of your own you edit `run.py`. The tree is short enough to read in one sitting: `checkpoint.py` and `model.py` hold the loading and the model definition, `runners.py` holds the execution, and `tokenizer.model` sits at the root so tokenization is bundled rather than downloaded. That is a good property for a reference release, because the path from raw weights to generated text has no hidden service in the middle.

The MoE layer was written to avoid custom kernels, and it says so

The README is unusually direct about its own performance: the implementation of the MoE layer in this repository is not efficient, and the choice was made to avoid the need for custom kernels so the correctness of the model could be validated. That is a reasonable engineering trade for a weights release, and it changes what you may conclude from it. Any throughput or latency you measure with this code describes JAX on your hardware and this reference implementation, not the model as it could be served. The same note applies to the quantization and sharding support listed in the specifications: activation sharding and 8-bit quantization are supported by the implementation, and the README does not document the flags that turn them on, so those live in `run.py` and `runners.py` where you have to read them. Treat this code as the validated reference and expect the serving work to be yours.

requirements.txt picks the CUDA 12 build of JAX for you

Four pinned packages, and one of them decides your hardware story.

text
dm_haiku==0.0.12
jax[cuda12-pip]==0.4.25 -f https://storage.googleapis.com/jax-releases/jax_cuda_releases.html
numpy==1.26.4
sentencepiece==0.2.0

The `cuda12-pip` extra plus the `-f` index URL means the install resolves a CUDA 12 build of JAX, so a machine without NVIDIA hardware, or one on a different accelerator stack, gets the wrong artifact and the repository documents no alternative. Every version is pinned with `==`, including `numpy==1.26.4`, so nothing moves on its own: a library fix or a compatibility change requires you to edit this file and own the result. Two details are worth knowing. `dm_haiku` is the neural network library the model is written against, and `sentencepiece` is what reads the bundled `tokenizer.model`. The 131,072 token vocabulary is part of the model, so changing that file changes the model rather than the plumbing.

The weights come by torrent or by the HuggingFace CLI, and only one is scripted

Two paths reach the checkpoint, and they have very different failure modes. The first is a magnet link for a torrent client, with public trackers listed inside the URI, which means you need a torrent client installed and reachable trackers to get 300-odd billion parameters onto a disk. The second is the HuggingFace Hub, and it is the path the README gives as a command:

bash
git clone https://github.com/xai-org/grok-1.git && cd grok-1
pip install huggingface_hub[hf_transfer]
huggingface-cli download xai-org/grok-1 --repo-type model --include ckpt-0/* --local-dir checkpoints --local-dir-use-symlinks False

`--include ckpt-0/*` fetches only the first checkpoint directory, and `--local-dir-use-symlinks False` materialises real files in `checkpoints` rather than pointers into a shared cache, which is what the loader in `checkpoint.py` expects. The expected layout is literal: a directory named `ckpt-0` inside `checkpoints`, and `checkpoints/` is present in the repository for exactly that reason. The `hf_transfer` extra exists because the download is large enough that the default transfer is the slow part.

Apache 2.0 covers the weights, and no release has ever been tagged

The licence position is unusually clear for a weights release. The code and the associated Grok-1 weights in this release are licensed under the Apache 2.0 licence, and the README states the scope precisely: the licence applies to the source files in this repository and the model weights of Grok-1, with the text in `LICENSE.txt`. Nothing else is covered by that grant, so read the terms for your own use rather than assuming a general permission. The maintenance picture is the other half. The repository is not archived, the default branch is `main`, there are no GitHub releases, and the last push was on 2024-08-30. There is no CI configuration among the top level entries either, so nothing runs this example on your behalf to tell you whether it still works, and `pyproject.toml` contains only a ruff configuration with no test target. The practical reading is that this is a frozen snapshot: usable as a reference, and yours to maintain from here.

Editorial conclusion

Use this repository when your goal is to confirm the Grok-1 weights load and generate, to read a compact reference for the architecture, or to start a serving stack you will write yourself. Do not adopt it as a running product: there is no server, no CLI surface, no test suite in the tree, and the MoE layer is explicitly not written for throughput. Before you build on it, check that your hardware matches a pinned `jax[cuda12-pip]==0.4.25` install, decide whether 8,192 tokens of context is enough for your use, and remember that the last push to this repository was on 2024-08-30 and that it publishes no releases, so a silent incompatibility with a current JAX version will be yours to find.

Frequently asked questions

is grok 1 open source

The weights and the code in this repository are released under the Apache 2.0 licence, and the README states the grant covers the source files here and the Grok-1 model weights. What is published is JAX example code for loading and sampling, not a trained service or an application.

how to use grok 1

Place the `ckpt-0` directory inside `checkpoints`, run `pip install -r requirements.txt` and then `python run.py`. The script loads the checkpoint and samples from the model on a test input, and the weights come either from the torrent magnet link or from the HuggingFace Hub with the `huggingface-cli download` command given in the README.

grok 1 vs gpt 4

This repository makes no comparison with GPT-4. What it specifies is the model itself: 314B parameters, a Mixture of 8 Experts with 2 used per token, 64 layers, 48 query and 8 key/value heads, a 131,072 token SentencePiece vocabulary and a maximum sequence length of 8,192 tokens.

Official sources

  1. Official README
  2. Project repository