darklife/darkriscv: a RISC-V softcore that fits in a thousand LUTs
opensouce RISC-V cpu core implemented in Verilog from scratch in one night!
At a glance
- What is it?
- A small Verilog CPU core written in one night, with a pipeline built to avoid stalls and forwards rather than handle them, aimed at FPGA and teaching use.
- Who is it for?
- darkriscv is most convincing as a study in how far a CPU design can be pushed by removing the mechanisms that usually handle hazards. The register file reads combinationally, the execute stage carries four ALUs in parallel, and nothing stalls or forwards, which is why the core reaches about one instruction per clock most of the time inside a very small LUT budget.
- Can I use it commercially?
- Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 32 days ago.
- What is it written in?
- Mainly Verilog, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 24, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Getting a simulation running with make and nothing else
The quick start is three commands, and the only prerequisite named is Icarus Verilog.
git clone [email protected]:darklife/darkriscv.git
cd darkriscv
makeRunning that produces a simulation of the whole SoC around the core, DarkSoCV, which boots default firmware, prints messages from the core itself, dumps pipeline information and writes a VCD trace. Opening the trace is a separate step:
gtkwave sim/darksocv.vcdThat is an unusually frictionless first run for a CPU project, and it matters for how you should read the repository. The top of the tree has `Makefile`, `BUILD.bazel`, `MODULE.bazel` and a Bazel version pin, so the same design is reachable from either GNU Make or Bazel. Directories are split by role rather than by pipeline stage: `rtl/` for the core logic, `sim/` for simulation and the VCD output, `src/` for the C and assembly firmware, `boards/` for board-specific constraints, and `doc/` for the block diagrams. There is also a `darkriscv.core` file at the root, which is the FuseSoC core description that lets the design be pulled into an existing hardware build system.
The pipeline has no interlocks, no stalls and no forwarding
The design decision that defines this core is stated plainly in the README: a DSP-like pipeline with no interlock, no stall and no forward between pipeline stages. Instead of detecting hazards and recovering, the design removes the dependency that would create them.
The instruction path has three stages, prefetch, instruction fetch and instruction decode, and then a single execute stage. Because execute is one stage, a result is available to the next instruction in the same cycle it is produced, so there is nothing to forward. The execute stage holds four ALUs working in parallel: a full ALU for register to register and register to immediate operations, a dedicated ALU for branch tests, a dedicated ALU for program counter update, and a dedicated ALU for memory address calculation.
The other half of the trick is the register bank. It is clocked and single path on write, but combinational and multi path on read, which is what allows four ALUs to be fed at once without arbitration. The cost is a wider register file and a fanout requirement that most ASIC flows would treat as expensive. On a small FPGA it is a reasonable trade, and the README quotes the result as roughly 70 percent of cycles at one instruction per clock.
RV32I and RV32E, with the extensions left optional
The baseline is the UCB RISC-V RV32I user space instruction set, with optional RV32E for the sixteen register variant. RV32E is not just a size saving here: the README notes it works better with LUT4 FPGAs, which is the smaller and older slice architecture still found on cheap boards.
Almost everything else in the feature list is a compile-time option rather than a default. That list includes CSRs for interrupts and debug, a sixteen by sixteen bit multiply accumulate instruction aimed at signal processing, a DBNZ instruction with delay slots for decrement and branch, coarse grained multithreading, machine level interrupt handling, supervisor level breakpoint handling, instruction and data caches, a Harvard to von Neumann bridge, an SDRAM controller taken from the kianRiscV project, big-endian support defined at build time, and the M extension limited to single cycle multiply with no divide.
Optional caches live in a block called DarkBridge rather than in the core, which is the right architectural choice. The core stays small and the memory system stays outside it, so a design that does not need a cache does not pay for one, and a design that needs one plugs it in at the bus. The Harvard interface is described as flexible for the same reason, easy to integrate a cache controller or a bus bridge.
What the core reaches on real FPGA parts
The performance claims are specific about parts, which makes them checkable in a way that vague megahertz numbers are not. The README lists up to 250MHz in an Ultrascale KU040, with 400MHz when overclocked, and up to 100MHz in a cheap Spartan-6. It also states the core fits in small Spartan-3E parts such as the XC3S100E, which is a useful lower bound for anyone budgeting LUTs.
Density is given in two forms: 233 DMIPS per thousand LUTs on a KU040 at 400MHz with IPC around 0.7, and 0.6 IPC per thousand LUTs at roughly 1200 LUTs. Area is quoted as between 850 and 1500 LUTs for the core alone in LUT6 technology, depending on which optional features and optimisations are enabled. Both figures are consistent with each other, and both are the kind of number that changes when someone turns on the MAC instruction or the machine mode CSRs, so read them as a range rather than a specification.
Supported boards are listed as Xilinx Spartan-3, Spartan-6, Spartan-7, Artix-7, Kintex-7 and Kintex Ultrascale, plus some real Altera and Lattice parts. On the toolchain side the README asks for GCC 9.0 or above for RISC-V, with GCC 12.0 needed for big-endian support.
DarkSoCV and the mixed Harvard and von Neumann arrangement
A softcore is not usable alone, so the project ships the surrounding SoC as DarkSoCV, and the diagram for it is more instructive than the core diagram. The arrangement described is mixed: the core runs with parallel Harvard caches for instruction and data, while the rest of the SoC works as a single von Neumann bus. Sequential instruction and data therefore share the same bus, which is what lets BRAM and SDRAM back both.
This is a practical compromise rather than a purity position. A true Harvard design would need separate memories for instruction and data, doubling the memory cost on a board where BRAM is the scarce resource. The mixed scheme keeps the core's instruction path clean, which is where the timing benefit comes from, while allowing one physical memory to serve both.
The history section explains where the design came from, and it is worth reading because it explains the small size. The core started as a proof of concept written between 2am and 8am on 19 August 2018, growing out of earlier sixteen bit RISC processors by the same author under the uDarkRISCV name. The first version was around three hundred lines of compact Verilog with a two stage pipeline. Two years of work produced the three stage pipeline with a single clock phase that shipped with GCC compiled code working on RV32I.
What is planned, and what a reader should conclude
The README separates shipped features from future work unusually clearly. The future list includes an Ethernet controller, symmetric multiprocessing, network on chip, RV64I support with the honest note that it is not as easy as it appears, dynamic bus sizing, user and supervisor modes, misaligned memory access, and a bridge for eight, sixteen and thirty-two bit buses.
Two things are worth noting about the repository itself. There are no tagged releases, so there is no version number to pin to, and the last push was on 2026-09-04. With five open issues, the visible queue is short. The absence of releases plus an optional-everything feature set means a reader has to read the source to know what a given build actually contains.
The honest summary is that this is an unusually clear teaching design rather than a product core. The pipeline strategy is the interesting part, and the documentation explains the reasoning well enough that the reasoning can be argued with. Where it stops short is integration: the cache, the bus bridge and the SDRAM controller are external, and a product would need to supply its own memory system and its own validation story.
Editorial conclusion
darkriscv is most convincing as a study in how far a CPU design can be pushed by removing the mechanisms that usually handle hazards. The register file reads combinationally, the execute stage carries four ALUs in parallel, and nothing stalls or forwards, which is why the core reaches about one instruction per clock most of the time inside a very small LUT budget. Start with `make` in the `sim/` directory and read the waveform, because the pipeline is far easier to understand from `sim/darksocv.vcd` than from prose. What the project does not settle is what you would use for a real product: the open issues are few, there are no tagged releases, and the feature list is full of optional switches and planned items rather than a settled design. For learning, demos and FPGA teaching, the source is unusually readable. For a system that has to run unattended, treat the surrounding SoC and toolchain as the part you still need to build.
Frequently asked questions
Which RISC-V CPU is the best?
For FPGA work and teaching, darkriscv is a reasonable pick because it fits in roughly 850 to 1500 LUTs and reaches about one instruction per clock most of the time. For something else, PicoRV32 and VexRiscv are the more common alternatives, and the SERP results for this project surface both. The best choice depends more on whether you need caches, multiprocessing or a documented bus interface than on instruction count.
Does darkriscv need an external cache and memory controller?
Yes. The core is a bare softcore and the cache, the Harvard to von Neumann bridge and the SDRAM controller live outside it in the DarkBridge block. The shipped DarkSoCV SoC supplies the surrounding system, which boots firmware and writes a VCD trace when you run the simulation.
Is darkriscv RV32I or RV32E by default?
The baseline is RV32I user space, and RV32E is an optional build configuration that gives the core sixteen registers instead of thirty two. The README notes RV32E works better on LUT4 FPGAs, which is where you are working with the smaller older slice architecture.
Can darkriscv run big-endian?
Yes, as a build time option. The README asks for GCC 12.0 or above when big-endian support is enabled, versus GCC 9.0 as the general baseline. Dynamic bus sizing and big-endian together are listed under future work, so the current support assumes a fixed bus width.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/darklife-darkriscv)