# vnmakarov/mir: a lightweight JIT compiler and C11 interpreter built on MIR

> MIR is a strongly typed intermediate representation with a JIT backend, an interpreter and a C11 compiler. It targets engineers embedding code generation inside a language runtime, not general application developers.

**vnmakarov/mir** — A lightweight JIT compiler based on MIR (Medium Internal Representation) and C11 JIT compiler and interpreter based on MIR

- Repository: https://github.com/vnmakarov/mir
- Stars: 2,661 · Forks: 181
- Language: C
- License: MIT
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/vnmakarov-mir

## The gap MIR fills between a bytecode loop and a full compiler framework

Most people who need dynamic code generation face the same fork in the road. They can write a threaded interpreter, which is simple but leaves performance on the table, or they can pull in a full compiler framework and inherit a large dependency, a build system, and an optimization pipeline they did not ask for. MIR occupies the space between those two. The README states the goal plainly: "MIR project goal is to provide a basis to implement fast and lightweight JITs." The repository is C, the license is MIT, and the tree ships the IR, the generator, the interpreter and the C front end together rather than as separate products.

The intended audience is narrow and specific. The README says the project plans to try the lightweight JIT first for CRuby or MRuby. That is a useful signal about the target user: someone maintaining a dynamic language implementation who wants to replace or supplement an interpreter loop. It is not a tool you install to make an existing Python or JavaScript program faster. MIR is infrastructure you link into your own runtime, and the value shows up only after you have written the code that emits MIR instructions.

That framing also explains the shape of the repository. The top level contains mir.c and mir.h for the IR itself, mir-gen.c plus one backend file per architecture, mir-interp.c for interpretation, and c2mir/ for the C11 front end. Nothing here is a command line product with a man page. The unit of adoption is a C library call, and the integration work is measured in the instructions your front end has to emit, not in configuration.

## How MIR represents code and where each backend takes over

MIR is a strongly typed IR, and the typing is not decorative. A module holds functions plus declarations and data. A function has a signature, local variables (arguments included) and instructions. Every local variable carries a type, and the README restricts that type to a 64-bit integer, float, double or long double. There is no 8-bit or 16-bit local. Narrow values appear only in memory operands, where the type can be an 8-, 16-, 32- or 64-bit signed or unsigned integer, and the README notes that an integer memory value is expanded with sign or zero promotion to 64 bits before use. If you are porting a front end that tracks narrow integer types precisely, that promotion rule is the first thing you have to reconcile.

Instructions are opcode plus operands, and the operand kinds are local variable, immediate, memory, label and reference. A memory operand is a small addressing expression: type, displacement, base and index local variables, and an integer constant scale for the index. References point at functions and declarations in the current module, in other MIR modules, or at C external functions. Control flow is labels plus branch instructions, with a combined compare-and-branch form, a switch instruction that jumps to one of several labels based on an index, and an indirect jump whose operand holds a label address previously taken.

The architecture split is visible in the file names. mir-gen-x86_64.c, mir-gen-aarch64.c, mir-gen-ppc64.c, mir-gen-s390x.c and mir-gen-riscv64.c are the code generators, and mir-gen-stub.c is the placeholder for a target without one. Two instruction groups exist specifically for dynamic language work. The README describes specialized light-weight call and return instructions for fast switching between a threaded interpreter and JITted code, and property instructions that generate specialized machine code when lazy basic block versioning is used. Those two features are the reason a Ruby-style runtime would look at MIR rather than a general purpose backend. Everything else in the IR is conventional: arithmetic, logical, comparison and conversion instructions over the 32- and 64-bit integer and floating point types, plus function and procedural call instructions and return instructions that optionally carry a value back.

## Installing MIR and running a first C program through c2mir

The repository has no package release and no published binary, so installation means building from source. INSTALL.md is the file the project provides for this, and the tree carries both a GNUmakefile and a CMakeLists.txt, so a make-based build or a CMake build are both available paths. The README does not document a package manager install, and the repository lists no homepage, so clone the source and follow INSTALL.md rather than looking for a distribution package.

Once built, the C front end is the quickest way to see MIR work. The README points to a Red Hat blog post titled "The MIR C interpreter and Just-in-Time (JIT) compiler" for the C2MIR compiler description, and c2mir/ is the directory that holds it. The c-tests/ and c-benchmarks/ directories contain the C inputs the project tests against, which is the practical reference for what the front end is expected to accept.

A minimal loop looks like this: build the tree, then compile a C file and run the result.

```bash
make
./c2mir -c -o out.mir input.c
```

Treat the flag spelling as something to confirm against INSTALL.md and the c2mir sources before you script it, because the README does not print a c2mir invocation and the exact options are the ones the tool's own usage message defines. What you should see after a successful run is a MIR file on disk, which you can then hand to the interpreter or to the JIT driver built from the same tree.

If you would rather skip the C front end and write MIR directly, the README's sieve example shows the textual form: a module declaration, an exported function, a func line with a return type and a named argument, local declarations, an alloca, and then instructions with labels.

```mir
m_sieve:  module
          export sieve
sieve:    func i32, i32:N
          local i64:iter, i64:count, i64:i, i64:k, i64:prime, i64:temp, i64:flags
          alloca flags, 819000
          mov iter, 0
loop:     bge
```

That text form is the fastest way to learn the operand model before writing API calls that construct the same module programmatically. The API path uses functions for creating modules, functions, instructions and operands, and the README notes you can also load MIR from a binary or text file, which means the text form and the API are two routes to the same in-memory structure.

## The disclaimer is the real support boundary

The README opens with a disclaimer, and it is unusually blunt: "There is absolutely no warranty that the code will work for any tests except ones given here and on platforms other than x86_64 Linux/OSX, aarch64 Linux/OSX(Apple M1), and ppc64le/s390x/riscv64 Linux." Read that carefully, because the phrasing is inverted from what you might expect. The tested set is x86_64 Linux and OSX, aarch64 Linux and OSX on Apple M1, and ppc64le, s390x and riscv64 on Linux. Everything else is outside the promise.

The topics list includes aarch64, apple, linux, macos, m1, ppc64, riscv64, s390x and x86-64, and the README embeds CI badges for AMD64 Linux/OSX/Windows, Apple Silicon, aarch64, ppc64le, s390x, riscv64 and an AMD64 Linux benchmark job. The Windows badge is worth noting next to the disclaimer, since Windows is not in the list of platforms the disclaimer names. A CI job running is not the same as a support commitment, and the disclaimer is the stronger statement of the two.

There is a second limitation that matters more for day-to-day work than any architecture list. The README does not document rollback, version compatibility between MIR binary files, or a stability guarantee for the IR across releases. MIR has a version number, currently 1.0.0 from the 2024-05-27 release, but the README does not state what a major or minor bump means for the binary format or the API. If you serialize MIR and store it, that silence is a risk you are taking on yourself, and the mitigation has to come from your own test suite rather than from a compatibility promise.

## Where MIR is the wrong choice

MIR is the wrong tool if your goal is to speed up an existing program without writing compiler code. There is no drop-in mode. You get a library, an IR, and a set of backend files, and the work of translating your program into MIR is yours. A project that wants a faster Python loop should look at the runtime's own JIT efforts, not at MIR.

It is also a poor fit when you need a documented API contract with a deprecation policy and long-term support. The README does not promise one, and the disclaimer explicitly withdraws warranty outside a short platform list. Teams that need to justify a dependency to a review board will find little to point at beyond the MIT license and the test directories.

Finally, MIR assumes a certain kind of input. The local variable types are 64-bit integer, float, double and long double. A front end that models a tagged integer narrower than 64 bits, or a garbage-collected pointer with precise write barriers, has to encode that discipline itself in the emitted instructions. The IR gives you property instructions and light-weight call and return instructions as building blocks for specialization, but the README describes them as mechanisms, not as a finished dynamic language runtime. The gap between "the IR can express this" and "the runtime does this" is where most of the engineering time goes, and it is not work the project does for you.

## MIR against SLJIT, GNU lightning and the LLVM-based backends

The natural comparisons are the other small JIT backends. SLJIT and GNU lightning both start from a similar premise: a compact library that emits machine code for a handful of architectures without an optimization pipeline. The difference with MIR is the level of abstraction. SLJIT and GNU lightning expose an assembler-like interface where you emit operations directly, while MIR asks you to build a typed IR with modules, functions, locals and labels, then hands that IR to a generator. MIR's IR is also the thing the interpreter consumes, so the same representation can be run without code generation at all. That is a meaningful architectural difference: an interpreter-only build is a supported configuration in the same tree, via mir-interp.c.

Against QBE, the split is similar but the input language differs. QBE takes its own SSA-based intermediate language and generates assembly, and it is designed around a text interface. MIR's IR is not described as SSA in the README, and MIR's front end is C rather than a dedicated IL. If you already emit QBE IL, moving to MIR means rewriting the emitter, not reusing it.

Against Cranelift and the LLVM-based JITs, the difference is weight and scope. Those projects bring optimization passes and a much larger codebase. MIR's stated goal is to be lightweight. That is a trade: you give up optimization machinery and, in exchange, get something you can read and build quickly. The README includes a benchmark CI badge, which suggests the project tracks performance, but it publishes no comparison numbers, so any claim about how MIR performs against these alternatives is not something the repository states.

## Maintenance, licensing and what a MIR upgrade costs

The repository is not archived, and the last push was on 2026-06-19. The CI badges cover AMD64 Linux/OSX/Windows, Apple Silicon, aarch64, ppc64le, s390x and riscv64, which indicates the test matrix is still being exercised. The release cadence is slower than the commit cadence: v0.1.1 on 2021-12-20, v0.1.2 on 2023-01-13, and v1.0.0 on 2024-05-27. No release after v1.0.0 appears in the repository's release list. A team that needs frequent tagged releases with changelogs should weigh that cadence against their own update policy.

The license is MIT, which is permissive and short. The practical implication for adopters is that you can link MIR into a proprietary product and ship it, provided you keep the copyright notice and permission notice with the distribution. This is not legal advice; read the LICENSE file in the repository and get your own counsel for anything that matters commercially.

Upgrade cost is the part the repository answers least well. The README does not describe a compatibility policy for the MIR binary format or the C API, and it does not document migration steps between versions. If you embed MIR, the safe assumption is that an upgrade may require re-reading MIR.md and rebuilding your emitter, and that stored MIR binary files should be regenerated rather than relied on. Verify that against the release notes for v1.0.0 before you plan a long-lived deployment, and treat the version number as a label rather than as a promise about format stability.

## Conclusion

Adopt MIR if you are building a language runtime and need a JIT backend plus a C11 front end you can read end to end, and you can accept that the README's own disclaimer limits tested support to x86_64 Linux/OSX, aarch64 Linux/OSX on Apple M1, and ppc64le/s390x/riscv64 Linux. Do not adopt it if you need a documented stability contract or a support channel; the README offers neither, and the LICENSE file is the only formal commitment in the repository. Before committing, check whether your target is on that tested list, read MIR.md for the instruction set your backend must emit, and confirm in c2mir/ whether the C11 features your input uses are covered.

## FAQ

### What is the purpose of a JIT compiler in the MIR project?

MIR's stated goal is to provide a basis for implementing fast and lightweight JITs, so the purpose here is to give a language runtime a code generator it can drive from its own front end. The same IR can also be run by the interpreter in mir-interp.c without generating machine code.

### What is the difference between an interpreter and a JIT compiler in MIR?

MIR ships both paths over the same representation: mir-interp.c executes MIR directly, while the mir-gen files generate machine code for x86_64, aarch64, ppc64, s390x and riscv64. The README also describes specialized light-weight call and return instructions for switching quickly between a threaded interpreter and JITted code.

### How does the MIR JIT actually work?

You build a module containing functions, typed local variables and instructions, then hand that IR to the generator, which has a separate backend file per architecture. Operands are local variables, immediates, memory references, labels or references to functions and declarations, and control flow uses labels with branch, switch and indirect jump instructions.

## Sources

- [Issues](https://github.com/vnmakarov/mir/issues)
- [License: MIT](https://github.com/vnmakarov/mir/blob/master/LICENSE)
- [README](https://github.com/vnmakarov/mir/blob/master/README.md)
- [Releases](https://github.com/vnmakarov/mir/releases)
- [vnmakarov/mir on GitHub](https://github.com/vnmakarov/mir)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/vnmakarov-mir
