microGPT-C: A Dependency-Free GPT in a Single C File
The most atomic way to train and inference a GPT in pure, dependency-free C
At a glance
- What is it?
- microGPT-C implements a complete character-level transformer, including forward pass, backpropagation, Adam optimizer, and sampling, in one C file with no dependencies beyond the C standard library. It targets developers and students who want to read, run, and modify a working GPT without a Python environment, a framework, or a GPU.
- Who is it for?
- microGPT-C suits anyone who wants to read a complete, working GPT implementation without installing Python or PyTorch. The single-file format and 4,192-parameter character-level model make the code approachable in a single sitting.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 45 days ago.
- What is it written in?
- Mainly C, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What microGPT-C Is and Who It Is For
microGPT-C is a character-level GPT implementation contained in a single C source file, src/microgpt.c, with no external dependencies beyond libc. It includes a complete forward pass with stored activations for backpropagation, a separate optimized single-token inference path, the Adam optimizer, and multinomial sampling. The model has 4,192 parameters.
The intended audience is developers or students who want to understand how a transformer learns from data without the overhead of a Python ecosystem. Because the entire model fits in one file, the code can be audited in a single reading session. It is not designed for conversational language modelling or for any corpus larger than what a laptop CPU can handle in seconds.
Architecture: Two Forward Passes and What Separates Them
The repository distinguishes between two forward-pass functions. gpt_forward stores intermediate activations so that the backward pass can compute gradients. gpt_forward_infer is a narrower path designed for single-token generation; it skips activation storage and runs faster, producing logits that match gpt_forward to within float32 rounding error.
Training and inference are therefore on separate code paths even though the model weights are shared. The README notes that docs/PERFORMANCE.md covers the mechanics of the inference path and the limits it imposes. This split matters practically: the inference path can reach tens of millions of tokens per second on modern hardware, while the training path prioritizes correctness and simplicity over raw speed.
The model trains on the names.txt corpus of roughly 32,000 names. Trained on 20,000 of them, it scores 2.2054 nats per character on those samples and 2.2039 on the 12,033 held-out names it never saw during training. According to the README, this beats an interpolated trigram baseline that uses nearly five times as many parameters, which indicates the model generalises rather than memorises.
Building and Running on Your Machine
The Makefile detects the host architecture and sets the appropriate flags. On ARM64 it passes -march=native, and on x86-64 it adds -mavx2 and -mfma alongside -O3 and -ffast-math. The build links only -lm.
To build and run the included names corpus in one step:
make runTo train on a different corpus where each item is on its own line:
./microgpt data/names.txtThe README shows what to expect during training: step counters printed every 5,000 steps with the current loss and a running average, followed by ten sampled names once training reaches step 20,000. On an Apple M5 Pro with NEON the inference path reaches 10,168,430 tokens per second; on an AMD Ryzen 5 5600H with AVX2 it reaches 6,927,775 tokens per second. The Makefile also supports macOS, Linux, and Windows via MSYS2.
Constraints: What the Design Rules Out
The 4,192-parameter model size is appropriate for character-level name generation but not for word-level or subword-level modelling at any practical scale. There is no tokenizer beyond single characters, so the architecture cannot handle vocabulary sizes typical of modern language models.
Because everything runs on the CPU, training on corpora significantly larger than the bundled names.txt will be slow. The optimized inference path handles single-token autoregressive decoding, but the README does not document a batched inference mode. Any use case requiring batch generation, long-context attention, or multi-GPU acceleration would need a different codebase.
The project also has no versioned releases. The last push to the repository was on 2026-08-17. There is no package distribution; users must clone the repository and compile from source using the provided Makefile. The clean target removes the binary and any generated object files.
microGPT-C Versus Python-Based Minimal GPTs
The closest comparable project is Andrej Karpathy's nanoGPT, which implements a byte-level or subword GPT in Python using PyTorch. nanoGPT supports GPU training, larger models, and existing tokenizers. The trade-off is a non-trivial dependency chain: Python, PyTorch, and often CUDA.
microGPT-C occupies a different position. It requires only a C compiler and libc. On a system without Python or pip, it compiles and trains in seconds. The single-file format is the point: a reader can step through the entire model from data loading to sampling without navigating a package tree or a virtual environment. Anyone already comfortable with Python who wants to scale up should use nanoGPT or one of the HuggingFace training frameworks. microGPT-C is for the reader who wants the smallest working example they can hold in their head at once.
License and Maintenance
microGPT-C is released under the MIT license, which permits use, modification, and redistribution without restriction provided the copyright notice is retained. The repository has no formal GitHub releases. The last push was on 2026-08-17.
The codebase is intentionally small. docs/PERFORMANCE.md provides documentation for the inference path beyond what the README covers. The Makefile includes a clean target for removing the compiled binary and object files. There is no test suite, no continuous integration configuration, and no package published to a registry. Because the implementation is a single C file, distributing a fork or a modified version requires no special build tooling beyond a standards-compliant C compiler.
Editorial conclusion
microGPT-C suits anyone who wants to read a complete, working GPT implementation without installing Python or PyTorch. The single-file format and 4,192-parameter character-level model make the code approachable in a single sitting. It is not the right tool for production language modelling, fine-tuning larger models, or any task requiring subword tokenization. Before using it as a learning reference, confirm that your C compiler supports the NEON or AVX2 intrinsics the Makefile enables, and read docs/PERFORMANCE.md for the documented constraints of the optimized single-token inference path.
Frequently asked questions
Does microGPT-C work on Windows?
The README states that it builds on Windows via MSYS2. The Makefile detects Windows via the OS environment variable and sets UNAME_M to x86_64, enabling AVX2 and FMA flags.
What corpus does microGPT-C train on by default?
The default corpus is data/names.txt, which contains approximately 32,000 names. The model trains on 20,000 of them and evaluates on the remaining 12,033 held-out names.
What is the difference between microGPT-C's two forward-pass functions?
gpt_forward stores intermediate activations needed by the backward pass during training. gpt_forward_infer is a faster single-token path used during inference that skips activation storage; the README states its logits match gpt_forward to within float32 rounding.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/vixhal-baraiya-microgpt-c)