gem5: a modular computer-system architecture simulator you build from source
The official repository for the gem5 computer-system architecture simulator.
At a glance
- What is it?
- gem5 is the C++ simulator behind a large share of academic architecture work, and it ships as source code you compile with SCons rather than as a ready-made binary. Here is what it does, how a first build and run go, and where it stops being the right tool.
- Who is it for?
- Adopt gem5 if you are doing architecture research, teaching, or system-software evaluation that needs cycle-level or system-level modelling of ARM, RISC-V, x86, POWER, MIPS or SPARC targets, and you can accept a source build and a configuration script you write yourself. Do not adopt it if you want a packaged binary, a GUI, or a drop-in emulator for running an unmodified workload at native speed; gem5 is a model of a machine, not a virtual machine for production use.
- Can I use it commercially?
- Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 7 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 24, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What gem5 models, and who actually needs that
gem5 is a simulator of computer systems, not a program that runs your code faster. The README describes it as a modular platform for computer-system architecture research, covering system-level architecture and processor microarchitecture, and names three uses: evaluating new hardware designs, system software changes, and compile-time and run-time optimizations. That list is the honest scope. If your question is "what would this cache hierarchy do to my kernel's throughput" or "how does this ISA extension change retirement behaviour," gem5 is built for it. If your question is "does my application work on ARM," a real board or an emulator answers it sooner.
The audience is narrower than the download numbers suggest. Architecture researchers, compiler and OS developers who need a controllable machine, and instructors who want students to see how a pipeline, cache or memory controller behaves under load. The cost is that you describe the machine yourself, in Python, before you can run anything. There is no default "just simulate my program" path in the repository; the configs directory holds example scripts, and the learning material on the website walks through writing your own.
The two-layer design: C++ objects driven by Python configuration
The architecture visible in the source tree is a split between a compiled simulator and a configuration layer. The src directory holds the C++ source, the Python wrappers, and the Python standard library, and the README states this explicitly. The C++ side implements the simulated components (cores, caches, interconnects, memory devices, and the ISA-specific decoding). The Python side instantiates them, wires their ports together, and sets their parameters, then hands the whole object graph to the compiled binary at startup.
That split explains the build options and the runtime interface. Because the ISAs are compiled in, a binary built for ARM cannot simulate x86 unless you built with ALL, and the README notes you can replace ALL with a single ISA name from ARM, NULL, MIPS, POWER, RISCV, SPARC or X86. Because configuration is Python, changing the number of cores or the cache size does not require a rebuild, only a different script. It also explains a common failure mode: a configuration script written for one release can break against another, because the parameters exposed to Python come from the C++ objects and change with them. The repository keeps configs, src and tests in the same tree for that reason.
Full-system simulation adds a third input. The README says that to run full-system simulations you may need compiled system firmware, kernel binaries and one or more disk images, depending on the configuration and workload, and points at resources.gem5.org for them. Those are separate downloads, not part of the build.
Building gem5 from source with SCons
gem5 is distributed as source. The README lists the required software: g++ or clang, Python (the simulator links in the Python interpreter), SCons, zlib, m4, and protobuf if you want trace capture and playback. Minimum versions are on the website's building page rather than in the README, so check there before you start; a mismatched compiler or Python is the most common reason a first build fails.
Once the dependencies are in place, the build is a single SCons invocation. This builds the optimized binary with every ISA compiled in:
scons build/ALL/gem5.optThe README says this produces an optimized version of the gem5 binary, gem5.opt, containing all gem5 ISAs. If you only need one, substitute the ISA name, for example X86 or ARM, in place of ALL. The complete list of names lives in the build_opts directory, which holds pre-made default configurations. Expect a long compile; this is a large C++ codebase and the build is the main time cost of getting started.
The result lands under build/ALL/ (or build/<ISA>/), and you run it by passing a configuration script as an argument. The README points to configs as the place where example simulation configuration scripts live, and the website's learning material covers writing your own. A first real use is to run one of those example scripts against a workload you supply, then read the statistics output the simulator writes at the end of the run. That statistics file, not the console log, is where the measurements are.
Where gem5 is the wrong tool
The first limitation is that gem5 does not give you a usable program output for free in full-system mode. You supply firmware, a kernel and a disk image, and you are responsible for getting your workload onto that image. The README treats these as external resources rather than bundled assets, which means a fresh clone is not a runnable system. Budget time for resource acquisition and image preparation before you budget time for experiments.
The second is speed. gem5 is a model; simulating a full operating system boot and a real workload takes orders of magnitude longer than running it natively, and that gap is the point of the tool rather than a defect. If your goal is functional testing, continuous integration of a software project, or anything where wall-clock time matters more than fidelity, an emulator or real hardware is the correct choice. gem5 will still "work" there, which is exactly why people waste weeks on it.
The third is configuration drift. Because the Python layer is generated from and bound to the C++ objects, a script that ran on one release may not run on the next without edits. The repository's own release history shows frequent tagged releases, and the stable branch receives pushes regularly (the last push was on 2026-09-23). If you need numbers you can reproduce months later, pin to a tag and keep your configuration scripts with your results.
Finally, the documentation is split. The README is a pointer document: it tells you to go to gem5.org for building, getting started and resources, and it does not document rollback, version compatibility between configuration scripts and binaries, or the statistics format. Those answers live on the website and in the source.
gem5 versus QEMU: a model against an emulator
QEMU is the comparison people reach for first, and the difference is in what each one is for. QEMU emulates a machine well enough to run unmodified software on it, and it is designed to do that quickly, including with dynamic translation and hardware acceleration. It answers "does this run on that architecture." gem5 simulates the timing and structure of a machine: the README frames it as a platform for architecture research covering system-level architecture and processor microarchitecture, which means the interesting output is a measurement (latency, occupancy, bandwidth, IPC) rather than a program result.
That distinction decides most tool choices. If you want to boot a Linux image, run a test suite, and get a pass or fail, QEMU is the shorter path and gem5 will feel like a detour with a long build. If you want to know how a change to the memory controller or the branch predictor affects execution, QEMU cannot tell you and gem5 is built to. Some workflows use both: QEMU for functional bring-up, gem5 for the measurements. Note also that gem5's own resource site distributes kernels and disk images, so the two tools can share workload artifacts even though their internal models are unrelated.
Licence, releases and the cost of staying current
The repository is BSD-3-Clause, per the LICENSE file at the top level. That is a permissive licence, which matters if you intend to modify the simulator, embed it in a tool, or ship results derived from it; permissive terms generally allow that, but the exact obligations (retention of the copyright notice and disclaimer) are set by the licence text itself. Read COPYING and LICENSE in the tree rather than a summary, and get legal advice if you are redistributing modified binaries. This is not legal advice.
Upgrade cost is driven by the release cadence and the coupling described above. Tagged releases in the recent history include v25.1.0.1, v25.1.0.0 and v25.0.0.1, so a major line appears roughly annually with point releases between. Moving between them means rebuilding the simulator and re-validating your configuration scripts, and if you use full-system mode, re-checking that your kernel and disk images still match the modelled platform. The build itself is the recurring cost: a full ALL build is long, and a single-ISA build is the practical way to keep iteration time down on a workstation.
Contribution paths are documented rather than implied. The README lists GitHub Discussions, GitHub Issues, a Jira tracker, Slack, and the gem5-users and gem5-dev mailing lists, and points at the website's contributing page and CONTRIBUTING.md. There is also a MAINTAINERS.yaml at the top level, which is where to look before sending a patch to a specific component.
Editorial conclusion
Adopt gem5 if you are doing architecture research, teaching, or system-software evaluation that needs cycle-level or system-level modelling of ARM, RISC-V, x86, POWER, MIPS or SPARC targets, and you can accept a source build and a configuration script you write yourself. Do not adopt it if you want a packaged binary, a GUI, or a drop-in emulator for running an unmodified workload at native speed; gem5 is a model of a machine, not a virtual machine for production use. Before committing, verify three things: that your target ISA is in the build_opts list, that the resources you need (firmware, kernel, disk image) are available from resources.gem5.org for your configuration, and that the Python configuration layer matches the version you built, since configs and src move together in this repository. The last push to the repository was on 2026-09-23, so the tree is moving; pin to a release tag such as v25.1.0.1 rather than tracking the stable branch if you need reproducible numbers.
Frequently asked questions
What is gem5 used for?
The README describes it as a modular platform for computer-system architecture research, covering system-level architecture and processor microarchitecture, and lists three uses: evaluating new hardware designs, system software changes, and compile-time and run-time system optimizations.
What is the gem5 simulator?
It is a simulator of computer systems written primarily in C++, with a Python configuration layer in the same source tree. The repository holds the full source code plus the tests and regressions used to validate it.
How do I run gem5?
You build the binary first, for example with scons build/ALL/gem5.opt, then run the resulting gem5.opt with a simulation configuration script from the configs directory. Full-system runs also need firmware, a kernel and a disk image, which the README says can be obtained from resources.gem5.org.
Who uses gem5?
The README frames the audience as people evaluating new hardware designs, system software changes, and compile-time and run-time optimizations, which in practice means architecture researchers and developers working on systems software. The repository does not publish user statistics.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/gem5-gem5)